Reinforcement learning
Rewards, policies, environments and results, for language models and beyond.
79 links, newest first.
- Reinforcement learningArticle
Exploration regimes for scaling LLM reinforcement learning
A CMU blog post describes three LLM RL exploration regimes: sharpening, chaining, and guided exploration. It says standard RL uses the first two and can plateau on hard problems, while mixing easy and hard data can cause interference.
The exploration and data-mixing challenges described can inform how engineers design RL training for difficult tasks.
- Reinforcement learningPost on X
A 14B Coding Model Fine-Tuned with SFT and DPO
The article describes fine-tuning a 14B model on coding conversations using an SFT and DPO pipeline, and reports numbers, costs, and frustrations.
It offers a concrete account of personal-data fine-tuning and the costs and results engineers may need to evaluate.
- Reinforcement learningRepository
slime v0.2.2 adds Int4-QAT and R3 support
The slime v0.2.2 release lists memory and performance improvements, Int4-QAT training, and full R3 support with DeepEP and MTP. It also upgrades to SGLang v0.5.7 and the Megatron dev branch.
Engineers using slime can review new training capabilities, rollout support, and performance changes.
- Reinforcement learningPost on X
Architecture choices affect continual learning and forgetting
The post describes a paper comparing architectures for continual learning. It says wider networks and removing or reducing global average pooling can improve retention, while ResNets and WideResNets learn new tasks quickly but forget more.
Architecture choices can influence the stability–plasticity trade-off, not just the continual-learning algorithm.
Self-Aligned Reward for LLM Reinforcement Learning
Self-Aligned Reward (SAR) is presented as a reward signal for LLM reinforcement-learning pipelines, compatible with PPO and GRPO. The post says it aims to improve reasoning quality while reducing token length.
Engineers can evaluate SAR as a reward module for balancing reasoning quality and generation length.
- Reinforcement learningArticle
Unsloth documents long-context GRPO training
Unsloth’s documentation describes long-context reinforcement learning fine-tuning with GRPO. The post claims its batching algorithms enable 380K context for gpt-oss on a 192GB GPU.
Useful for engineers exploring long-context RL fine-tuning and its stated hardware requirements.
- Reinforcement learningPost on X
Bayesian Generalization for Offline Model-Based RL
The post describes a paper that explores dropping conservative constraints in offline RL and relying on Bayesian principles for adaptive generalization. It reports that long-horizon rollouts make this approach work.
It highlights an alternative to conservative offline RL that engineers can evaluate for model-based learning.
- Reinforcement learningPost on X
Interviews on the RL environment ecosystem
Chris Barber and JS Denain interviewed 18 people from RL environment startups, neolabs, and frontier labs about a field the post says mostly stays behind closed doors.
The interviews offer engineers a view into how RL environments are being developed across organizations.
Comparing SFT and reinforcement fine-tuning for continual post-training
The paper compares supervised fine-tuning and reinforcement fine-tuning for continual post-training of foundation models. Its title states that reinforcement fine-tuning naturally mitigates forgetting.
The comparison may help engineers choose post-training approaches when adapting models to evolving tasks.
Full-stack fine-tuning for the Q programming language
The technical report presents an open-source approach for adapting language models to Q, a programming language used in quantitative finance. It addresses the challenge of applying LLMs to niche languages under-represented in training data.
The report may offer engineers a practical approach to adapting LLMs for specialized programming languages.
- Reinforcement learningPost on X
SETA Releases RL Environments for Terminal Agents
SETA says it released 400 terminal-agent training environments, an RL training pipeline, and weights for SETA-RL-Qwen3-8B. It also describes its CAMEL terminal toolkit harness as state of the art on terminal-bench.
The release could provide engineers with environments, a training pipeline, and model weights for terminal-agent RL experiments.
- Reinforcement learningRepository
OpenTinker provides RL-as-a-Service infrastructure for foundation models
OpenTinker is an open-source platform for agentic reinforcement learning as a service. It separates programming from execution and supports transition from training to inference.
The platform offers an alternative infrastructure for building and deploying reinforcement-learning workflows for foundation models.
- Reinforcement learningPost on X
Recursive language models use code to process long prompts
The post describes MIT’s recursive language models, which use a Python interface to search and filter long prompts and recursively process smaller chunks. It claims they handle inputs up to about 100× larger than usual and outperform base models and retrieval setups on information-heavy tasks.
The approach may help engineers process long inputs without relying on a larger context window.
- Reinforcement learningPost on X
Mask-Progressive RL Distillation for VLMs
The post describes Masters, a mask-progressive RL distillation framework for VLMs that it says avoids the computational cost of online RL. It reports scores rising from 75.7% to 80.4% for Qwen3-VL-8B and 75.4% to 80.0% for InternVL3.5-8B.
The reported results may interest engineers evaluating lower-compute approaches to improving VLM performance.
JustRL tests a simple RL recipe on 1.5B reasoning models
The paper presents JustRL, a single-stage training approach with fixed hyperparameters for two 1.5B reasoning models. Its preview reports average accuracies of 54.9% and 64.3%.
It offers a baseline for comparing simpler RL training recipes with more complex approaches.
- Reinforcement learningArticle
LoRA strategies for multi-domain RLVR
The author reports experiments on joint multi-domain RL training and expert LoRAs, finding interference with joint training. Averaging expert LoRA weights and continuing training outperformed a gated approach in their experiments.
The work compares practical PEFT strategies for RLVR and reports a simple baseline that beat the tested gating method.
- Reinforcement learningPost on X
Self-play RL trains coding agents on bug injection and repair
Self-play SWE-RL trains one LLM agent to inject and repair bugs in real repositories. The post says the agent receives no natural-language problem description or human-labeled issues.
The approach explores training software agents without relying on human-curated coding tasks.
Advantage shaping for single-rollout multimodal RL
The post says a single-rollout approach for multimodal RL became stable after adding advantage shaping with an entropy bonus. The linked paper preview describes Single-stream Policy Optimization for LLMs.
The stabilization detail may be useful when designing low-variance policy-gradient methods.
- Reinforcement learningRepository
OpenTinker separates RL programming from execution
OpenTinker is an open-source RL-as-a-Service platform for foundation models. Its page describes separating programming from execution and supporting a transition from training to inference.
The design may help engineers develop RL environments locally while using remote compute.
- Reinforcement learningArticle
A roundup of RL training systems and pass@k results
The post surveys on-policy rollout mismatch, rollout-system design, and studies of RL’s effects on pass@k. It reports mixed findings, including gains from relaxing reward assignment constraints and from OLMo 3 DPO training.
It highlights practical rollout challenges and evidence that RL’s effect on sampling performance depends on the training approach.
Build with AgentLog
Send your own newsletter
CommsHarbor keeps contacts, consent and one-click unsubscribe together.
Free workspace
Open CommsHarbor


