Training and fine-tuning
Pre-training, post-training, LoRA and the recipes behind better models.
108 links, newest first.
RLPT Applies Reinforcement Learning to Pre-Training Data
The post introduces Reinforcement Learning on Pre-training Data (RLPT), which uses next-segment rewards to train models on pre-training data without additional human annotations. It reports benchmark gains for Qwen3-4B-Base.
The approach offers an alternative way to scale reinforcement learning using pre-training data.
- Training and fine-tuningArticle
RF-DETR fine-tuned for basketball detection
The post describes starting with an RF-DETR model fine-tuned to detect players, numbers, referees, the ball, and the rim. It links to a basketball player detection dataset.
Useful as a concrete example of fine-tuning a detection model for basketball analysis.
- Training and fine-tuningPost on X
SWE-Swiss training recipe for software engineering models
Researchers from Peking University, ByteDance Seed, and HKU introduce SWE-Swiss, a training recipe for software engineering tasks. They report that the 32B-parameter SWE-Swiss-32B scores 60.2% on SWE-bench Verified.
The reported results may help engineers assess training approaches for software engineering models.
- Training and fine-tuningRepository
DeepSpeed Domino for distributed LLM training
The post describes DeepSpeed Domino as a tensor parallelism engine designed to reduce communication overhead during LLM training, with communication hiding and multi-node scaling.
Engineers evaluating distributed training optimizations can review Domino’s implementation and design.
- Training and fine-tuningPost on X
PEER uses product keys to retrieve from over a million experts
PEER is a Mixture of Experts architecture with product-key routing and single-neuron MLP experts. It retrieves a small set of experts for each input and combines their outputs with router scores.
The design offers engineers an approach to scaling model capacity through sparse expert retrieval.
- Training and fine-tuningPost on X
A PyTorch LoRA Fine-Tuning Example for a 5M-Parameter MLP
The author implemented LoRA to fine-tune a simple MLP for classification. The exercise freezes the original weights, trains about 9,000 adapter parameters on one label class, and compares performance with the adapter enabled and disabled.
It gives engineers a compact example of applying LoRA adapters and evaluating their effect.
- Training and fine-tuningPost on X
Fine-tuning Llama 3 8B with knowledge generated by Llama 3 405B
An open-source Colab notebook fine-tunes Llama 3 8B on domain-specific knowledge generated by Llama 3 405B.
Engineers can use the notebook as a starting point for distilling domain knowledge from a larger model into a smaller one.
ESFT fine-tunes task-relevant experts in sparse LLMs
The paper studies parameter-efficient fine-tuning for Mixture-of-Experts LLMs, where this approach remains underexplored. ESFT trains selected experts relevant to customized tasks.
Engineers working with MoE models can explore a fine-tuning approach that targets task-relevant experts.
Build with AgentLog
Send your own newsletter
CommsHarbor keeps contacts, consent and one-click unsubscribe together.
Free workspace
Open CommsHarbor