Skip to content
EN

Training and fine-tuning

Pre-training, post-training, LoRA and the recipes behind better models.

108 links, newest first.

Get the weekly briefing

The best new links of the topics you pick, summarized with the source. At most one email a week.

Topics: Training and fine-tuning

Before the first issue we email you to confirm; leaving takes one click. Sent with CommsHarbor. Privacy

  1. RLPT Applies Reinforcement Learning to Pre-Training Data

    The post introduces Reinforcement Learning on Pre-training Data (RLPT), which uses next-segment rewards to train models on pre-training data without additional human annotations. It reports benchmark gains for Qwen3-4B-Base.

    The approach offers an alternative way to scale reinforcement learning using pre-training data.

  2. RF-DETR fine-tuned for basketball detection

    The post describes starting with an RF-DETR model fine-tuned to detect players, numbers, referees, the ball, and the rim. It links to a basketball player detection dataset.

    Useful as a concrete example of fine-tuning a detection model for basketball analysis.

  3. SWE-Swiss training recipe for software engineering models

    Researchers from Peking University, ByteDance Seed, and HKU introduce SWE-Swiss, a training recipe for software engineering tasks. They report that the 32B-parameter SWE-Swiss-32B scores 60.2% on SWE-bench Verified.

    The reported results may help engineers assess training approaches for software engineering models.

  4. DeepSpeed Domino for distributed LLM training

    The post describes DeepSpeed Domino as a tensor parallelism engine designed to reduce communication overhead during LLM training, with communication hiding and multi-node scaling.

    Engineers evaluating distributed training optimizations can review Domino’s implementation and design.

  5. PEER uses product keys to retrieve from over a million experts

    PEER is a Mixture of Experts architecture with product-key routing and single-neuron MLP experts. It retrieves a small set of experts for each input and combines their outputs with router scores.

    The design offers engineers an approach to scaling model capacity through sparse expert retrieval.

  6. A PyTorch LoRA Fine-Tuning Example for a 5M-Parameter MLP

    The author implemented LoRA to fine-tune a simple MLP for classification. The exercise freezes the original weights, trains about 9,000 adapter parameters on one label class, and compares performance with the adapter enabled and disabled.

    It gives engineers a compact example of applying LoRA adapters and evaluating their effect.

  7. Fine-tuning Llama 3 8B with knowledge generated by Llama 3 405B

    An open-source Colab notebook fine-tunes Llama 3 8B on domain-specific knowledge generated by Llama 3 405B.

    Engineers can use the notebook as a starting point for distilling domain knowledge from a larger model into a smaller one.

  8. ESFT fine-tunes task-relevant experts in sparse LLMs

    The paper studies parameter-efficient fine-tuning for Mixture-of-Experts LLMs, where this approach remains underexplored. ESFT trains selected experts relevant to customized tasks.

    Engineers working with MoE models can explore a fine-tuning approach that targets task-relevant experts.

Build with AgentLog

List your MCP, skill or plugin

Reach the engineers who read these briefings.

Sponsor AgentLog

Footer, sidebar or featured slot for 30 days.

From US$ 60

See the slots

Send your own newsletter

CommsHarbor keeps contacts, consent and one-click unsubscribe together.

Free workspace

Open CommsHarbor