Skip to content
EN

Training and fine-tuning

Pre-training, post-training, LoRA and the recipes behind better models.

108 links, newest first.

Get the weekly briefing

The best new links of the topics you pick, summarized with the source. At most one email a week.

Topics: Training and fine-tuning

Before the first issue we email you to confirm; leaving takes one click. Sent with CommsHarbor. Privacy

  1. A guide to LoRA variants for parameter-efficient fine-tuning

    The Turing Post guide surveys more than 15 LoRA variants, including QLoRA, DoRA, and Mixture-of-LoRA-Experts, tracing the evolution of parameter-efficient fine-tuning for large language models.

    Useful for engineers comparing approaches to parameter-efficient fine-tuning.

  2. Ouroboros adds input-conditioned LoRA to recursive transformers

    Ouroboros attaches a compact controller to a shared transformer block, generating per-step weight modulation from the hidden state. On Qwen2.5-3B, the authors report lower training loss than a 17-layer baseline, while noting gains are in-distribution.

    It explores how a small controller can add step-specific transformations when a transformer block is reused.

  3. DeepSeek releases Tile Kernels for LLM operations

    DeepSeek released GPU kernels for LLM operations built with TileLang. The company says some kernels are used in internal training and inference.

    Engineers can evaluate GPU kernel optimizations relevant to LLM training and inference.

  4. Megatron-LM uses model parallelism to train large language models

    This 2019 NVIDIA paper describes model parallelism for training multi-billion-parameter language models by splitting model weights across GPUs.

    It explains a technique for training models whose weights do not fit on a single GPU.

  5. An explanation of rotary positional embeddings

    The post links to an article titled “You could have designed state of the art positional encoding” and describes it as an explanation of RoPE.

    Useful for engineers looking to understand positional encoding in model design.

  6. SHINE maps context to LoRA adapters in a single pass

    The paper presents SHINE, a hypernetwork that maps diverse contexts to LoRA adapters for LLMs. Its design reuses the frozen LLM’s parameters; the authors also introduce pretraining and instruction fine-tuning.

    Engineers can explore an approach to turning context into adapters without updating the base model.

  7. Unsloth’s notebooks for fine-tuning and RL

    The Unsloth repository collects more than 250 notebooks for fine-tuning and RL across text, vision, audio, embedding, and TTS models.

    Engineers can use the notebooks as examples for training workflows across several model types.

  8. Auditing LoRA as Knowledge Memory

    The post describes a paper that studies how LoRA memory storage scales and saturates with rank. It reports higher memory per token when training on QA or summaries rather than raw passages.

    The findings may help engineers choose training data and rank when using LoRA to store swappable knowledge.

  9. Chunking SIGReg slices reduces VRAM use in LeJEPA

    The author reports that splitting 256 SIGReg slices into chunks of 32 used 8× less VRAM and finished faster on a 4090 MaxQ. They identify large complex exponential tensors as the bottleneck.

    A practical memory and runtime optimization for engineers working with LeJEPA on constrained GPUs.

  10. An Empirical Analysis of LoRA as Knowledge Memory

    The paper studies LoRA as a modular, parametric memory for updating knowledge in pretrained LLMs. It maps the design space and examines capacity limits, multi-module scaling, and merging interference.

    It helps engineers assess when LoRA adapters can store knowledge reliably and where their limits may affect a system.

  11. DAPO reports improving GRPO AIME results from 30% to 50%

    The DAPO paper reports that standard GRPO achieved 30% on AIME in its experiments, while adding a couple of tricks raised the result to 50%.

    The reported comparison may help engineers evaluate techniques for improving reinforcement-learning training.

  12. Prime Intellect’s RL training recipes

    Prime Intellect’s documentation provides practical RL training recipes for math reasoning, code generation, tool use, and more on Lab.

    Engineers can use the recipes as practical starting points for RL training tasks.

  13. QLoRA Fine-Tuning Demo Uses 50 GPT-5.2 Outputs

    The post describes a local distillation demonstration: 50 outputs from GPT-5.2 are used to QLoRA fine-tune Llama 3B. The author claims this transfers substantial capability.

    It offers a concrete example of using teacher-model outputs for QLoRA fine-tuning.

  14. TurningPoint-GRPO Uses Step-Level Rewards for Flow Matching

    TurningPoint-GRPO is described as a framework for sparse rewards in flow matching models. It detects steps that reverse the local reward trend via sign changes to capture long-term dependencies.

    The approach may interest engineers exploring reward design for flow matching models.

  15. QED-Nano: Training a 4B Model to Prove Olympiad Problems

    A blog post about QED-Nano, described as teaching a tiny model to prove hard theorems. The post links to a collection and code.

    Engineers can inspect the project’s training details, datasets, and implementation.

  16. A PyTorch Notebook on Improving GRPO

    Chapter 7 of Reasoning From Scratch adds and analyzes clipped policy ratios, a KL term, format rewards, and other improvements to a GRPO implementation.

    Engineers can inspect a step-by-step implementation of GRPO improvements for reasoning models.

  17. Training-Free GRPO Uses Textual Experience Memory

    The post describes a method that keeps the model frozen and uses a natural-language memory of successful and failed rollouts. It claims the approach matches fine-tuned performance with 100 examples.

    The approach may interest engineers exploring ways to reduce the compute and parameter updates involved in model training.

  18. WeightWatch detects late-stage generalization collapse in grokking

    The paper identifies anti-grokking, a late-stage collapse of generalization, in two extended grokking experiments: a 3-layer MLP on a subset of MNIST and a transformer trained on modular addition. The post says WeightWatcher detects signatures in layer weight matrices without training or test data.

    It explores whether model weights can reveal a loss of generalization during prolonged training.

  19. EB-JEPA: An Open-Source Library for Joint-Embedding Predictive Architectures

    The paper presents EB-JEPA, an open-source library with modular implementations for learning representations and world models using JEPAs. These architectures predict in representation space rather than pixel space.

    Engineers can explore implementations of JEPA-based representation learning for downstream tasks.

  20. Book chapter introduces GRPO for verifiable-reward RL

    A chapter of *Build a Reasoning Model (From Scratch)* covers GRPO from scratch for reinforcement learning with verifiable rewards. The author says follow-up material will explore additional runs, analyses, and algorithmic tweaks.

    Useful for engineers seeking a step-by-step introduction to GRPO and reasoning-model training.

Build with AgentLog

List your MCP, skill or plugin

Reach the engineers who read these briefings.

Sponsor AgentLog

Footer, sidebar or featured slot for 30 days.

From US$ 60

See the slots

Send your own newsletter

CommsHarbor keeps contacts, consent and one-click unsubscribe together.

Free workspace

Open CommsHarbor