Skip to content
EN

Reinforcement learning

Rewards, policies, environments and results, for language models and beyond.

79 links, newest first.

Get the weekly briefing

The best new links of the topics you pick, summarized with the source. At most one email a week.

Topics: Reinforcement learning

Before the first issue we email you to confirm; leaving takes one click. Sent with CommsHarbor. Privacy

  1. Exploration regimes for scaling LLM reinforcement learning

    A CMU blog post describes three LLM RL exploration regimes: sharpening, chaining, and guided exploration. It says standard RL uses the first two and can plateau on hard problems, while mixing easy and hard data can cause interference.

    The exploration and data-mixing challenges described can inform how engineers design RL training for difficult tasks.

  2. A 14B Coding Model Fine-Tuned with SFT and DPO

    The article describes fine-tuning a 14B model on coding conversations using an SFT and DPO pipeline, and reports numbers, costs, and frustrations.

    It offers a concrete account of personal-data fine-tuning and the costs and results engineers may need to evaluate.

  3. slime v0.2.2 adds Int4-QAT and R3 support

    The slime v0.2.2 release lists memory and performance improvements, Int4-QAT training, and full R3 support with DeepEP and MTP. It also upgrades to SGLang v0.5.7 and the Megatron dev branch.

    Engineers using slime can review new training capabilities, rollout support, and performance changes.

  4. Architecture choices affect continual learning and forgetting

    The post describes a paper comparing architectures for continual learning. It says wider networks and removing or reducing global average pooling can improve retention, while ResNets and WideResNets learn new tasks quickly but forget more.

    Architecture choices can influence the stability–plasticity trade-off, not just the continual-learning algorithm.

  5. Self-Aligned Reward for LLM Reinforcement Learning

    Self-Aligned Reward (SAR) is presented as a reward signal for LLM reinforcement-learning pipelines, compatible with PPO and GRPO. The post says it aims to improve reasoning quality while reducing token length.

    Engineers can evaluate SAR as a reward module for balancing reasoning quality and generation length.

  6. Unsloth documents long-context GRPO training

    Unsloth’s documentation describes long-context reinforcement learning fine-tuning with GRPO. The post claims its batching algorithms enable 380K context for gpt-oss on a 192GB GPU.

    Useful for engineers exploring long-context RL fine-tuning and its stated hardware requirements.

  7. Bayesian Generalization for Offline Model-Based RL

    The post describes a paper that explores dropping conservative constraints in offline RL and relying on Bayesian principles for adaptive generalization. It reports that long-horizon rollouts make this approach work.

    It highlights an alternative to conservative offline RL that engineers can evaluate for model-based learning.

  8. Interviews on the RL environment ecosystem

    Chris Barber and JS Denain interviewed 18 people from RL environment startups, neolabs, and frontier labs about a field the post says mostly stays behind closed doors.

    The interviews offer engineers a view into how RL environments are being developed across organizations.

  9. Comparing SFT and reinforcement fine-tuning for continual post-training

    The paper compares supervised fine-tuning and reinforcement fine-tuning for continual post-training of foundation models. Its title states that reinforcement fine-tuning naturally mitigates forgetting.

    The comparison may help engineers choose post-training approaches when adapting models to evolving tasks.

  10. Full-stack fine-tuning for the Q programming language

    The technical report presents an open-source approach for adapting language models to Q, a programming language used in quantitative finance. It addresses the challenge of applying LLMs to niche languages under-represented in training data.

    The report may offer engineers a practical approach to adapting LLMs for specialized programming languages.

  11. SETA Releases RL Environments for Terminal Agents

    SETA says it released 400 terminal-agent training environments, an RL training pipeline, and weights for SETA-RL-Qwen3-8B. It also describes its CAMEL terminal toolkit harness as state of the art on terminal-bench.

    The release could provide engineers with environments, a training pipeline, and model weights for terminal-agent RL experiments.

  12. OpenTinker provides RL-as-a-Service infrastructure for foundation models

    OpenTinker is an open-source platform for agentic reinforcement learning as a service. It separates programming from execution and supports transition from training to inference.

    The platform offers an alternative infrastructure for building and deploying reinforcement-learning workflows for foundation models.

  13. Recursive language models use code to process long prompts

    The post describes MIT’s recursive language models, which use a Python interface to search and filter long prompts and recursively process smaller chunks. It claims they handle inputs up to about 100× larger than usual and outperform base models and retrieval setups on information-heavy tasks.

    The approach may help engineers process long inputs without relying on a larger context window.

  14. Mask-Progressive RL Distillation for VLMs

    The post describes Masters, a mask-progressive RL distillation framework for VLMs that it says avoids the computational cost of online RL. It reports scores rising from 75.7% to 80.4% for Qwen3-VL-8B and 75.4% to 80.0% for InternVL3.5-8B.

    The reported results may interest engineers evaluating lower-compute approaches to improving VLM performance.

  15. JustRL tests a simple RL recipe on 1.5B reasoning models

    The paper presents JustRL, a single-stage training approach with fixed hyperparameters for two 1.5B reasoning models. Its preview reports average accuracies of 54.9% and 64.3%.

    It offers a baseline for comparing simpler RL training recipes with more complex approaches.

  16. LoRA strategies for multi-domain RLVR

    The author reports experiments on joint multi-domain RL training and expert LoRAs, finding interference with joint training. Averaging expert LoRA weights and continuing training outperformed a gated approach in their experiments.

    The work compares practical PEFT strategies for RLVR and reports a simple baseline that beat the tested gating method.

  17. Self-play RL trains coding agents on bug injection and repair

    Self-play SWE-RL trains one LLM agent to inject and repair bugs in real repositories. The post says the agent receives no natural-language problem description or human-labeled issues.

    The approach explores training software agents without relying on human-curated coding tasks.

  18. Advantage shaping for single-rollout multimodal RL

    The post says a single-rollout approach for multimodal RL became stable after adding advantage shaping with an entropy bonus. The linked paper preview describes Single-stream Policy Optimization for LLMs.

    The stabilization detail may be useful when designing low-variance policy-gradient methods.

  19. OpenTinker separates RL programming from execution

    OpenTinker is an open-source RL-as-a-Service platform for foundation models. Its page describes separating programming from execution and supporting a transition from training to inference.

    The design may help engineers develop RL environments locally while using remote compute.

  20. A roundup of RL training systems and pass@k results

    The post surveys on-policy rollout mismatch, rollout-system design, and studies of RL’s effects on pass@k. It reports mixed findings, including gains from relaxing reward assignment constraints and from OLMo 3 DPO training.

    It highlights practical rollout challenges and evidence that RL’s effect on sampling performance depends on the training approach.

Build with AgentLog

List your MCP, skill or plugin

Reach the engineers who read these briefings.

Sponsor AgentLog

Footer, sidebar or featured slot for 30 days.

From US$ 60

See the slots

Send your own newsletter

CommsHarbor keeps contacts, consent and one-click unsubscribe together.

Free workspace

Open CommsHarbor