Training and fine-tuning
Pre-training, post-training, LoRA and the recipes behind better models.
108 links, newest first.
- Training and fine-tuningArticle
A guide to LoRA variants for parameter-efficient fine-tuning
The Turing Post guide surveys more than 15 LoRA variants, including QLoRA, DoRA, and Mixture-of-LoRA-Experts, tracing the evolution of parameter-efficient fine-tuning for large language models.
Useful for engineers comparing approaches to parameter-efficient fine-tuning.
Ouroboros adds input-conditioned LoRA to recursive transformers
Ouroboros attaches a compact controller to a shared transformer block, generating per-step weight modulation from the hidden state. On Qwen2.5-3B, the authors report lower training loss than a 17-layer baseline, while noting gains are in-distribution.
It explores how a small controller can add step-specific transformations when a transformer block is reused.
- Training and fine-tuningPost on X
DeepSeek releases Tile Kernels for LLM operations
DeepSeek released GPU kernels for LLM operations built with TileLang. The company says some kernels are used in internal training and inference.
Engineers can evaluate GPU kernel optimizations relevant to LLM training and inference.
Megatron-LM uses model parallelism to train large language models
This 2019 NVIDIA paper describes model parallelism for training multi-billion-parameter language models by splitting model weights across GPUs.
It explains a technique for training models whose weights do not fit on a single GPU.
- Training and fine-tuningArticle
An explanation of rotary positional embeddings
The post links to an article titled “You could have designed state of the art positional encoding” and describes it as an explanation of RoPE.
Useful for engineers looking to understand positional encoding in model design.
SHINE maps context to LoRA adapters in a single pass
The paper presents SHINE, a hypernetwork that maps diverse contexts to LoRA adapters for LLMs. Its design reuses the frozen LLM’s parameters; the authors also introduce pretraining and instruction fine-tuning.
Engineers can explore an approach to turning context into adapters without updating the base model.
- Training and fine-tuningRepository
Unsloth’s notebooks for fine-tuning and RL
The Unsloth repository collects more than 250 notebooks for fine-tuning and RL across text, vision, audio, embedding, and TTS models.
Engineers can use the notebooks as examples for training workflows across several model types.
- Training and fine-tuningPost on X
Auditing LoRA as Knowledge Memory
The post describes a paper that studies how LoRA memory storage scales and saturates with rank. It reports higher memory per token when training on QA or summaries rather than raw passages.
The findings may help engineers choose training data and rank when using LoRA to store swappable knowledge.
- Training and fine-tuningPost on X
Chunking SIGReg slices reduces VRAM use in LeJEPA
The author reports that splitting 256 SIGReg slices into chunks of 32 used 8× less VRAM and finished faster on a 4090 MaxQ. They identify large complex exponential tensors as the bottleneck.
A practical memory and runtime optimization for engineers working with LeJEPA on constrained GPUs.
An Empirical Analysis of LoRA as Knowledge Memory
The paper studies LoRA as a modular, parametric memory for updating knowledge in pretrained LLMs. It maps the design space and examines capacity limits, multi-module scaling, and merging interference.
It helps engineers assess when LoRA adapters can store knowledge reliably and where their limits may affect a system.
- Training and fine-tuningPost on X
DAPO reports improving GRPO AIME results from 30% to 50%
The DAPO paper reports that standard GRPO achieved 30% on AIME in its experiments, while adding a couple of tricks raised the result to 50%.
The reported comparison may help engineers evaluate techniques for improving reinforcement-learning training.
Prime Intellect’s RL training recipes
Prime Intellect’s documentation provides practical RL training recipes for math reasoning, code generation, tool use, and more on Lab.
Engineers can use the recipes as practical starting points for RL training tasks.
- Training and fine-tuningPost on X
QLoRA Fine-Tuning Demo Uses 50 GPT-5.2 Outputs
The post describes a local distillation demonstration: 50 outputs from GPT-5.2 are used to QLoRA fine-tune Llama 3B. The author claims this transfers substantial capability.
It offers a concrete example of using teacher-model outputs for QLoRA fine-tuning.
- Training and fine-tuningPost on X
TurningPoint-GRPO Uses Step-Level Rewards for Flow Matching
TurningPoint-GRPO is described as a framework for sparse rewards in flow matching models. It detects steps that reverse the local reward trend via sign changes to capture long-term dependencies.
The approach may interest engineers exploring reward design for flow matching models.
QED-Nano: Training a 4B Model to Prove Olympiad Problems
A blog post about QED-Nano, described as teaching a tiny model to prove hard theorems. The post links to a collection and code.
Engineers can inspect the project’s training details, datasets, and implementation.
- Training and fine-tuningRepository
A PyTorch Notebook on Improving GRPO
Chapter 7 of Reasoning From Scratch adds and analyzes clipped policy ratios, a KL term, format rewards, and other improvements to a GRPO implementation.
Engineers can inspect a step-by-step implementation of GRPO improvements for reasoning models.
- Training and fine-tuningPost on X
Training-Free GRPO Uses Textual Experience Memory
The post describes a method that keeps the model frozen and uses a natural-language memory of successful and failed rollouts. It claims the approach matches fine-tuned performance with 100 examples.
The approach may interest engineers exploring ways to reduce the compute and parameter updates involved in model training.
WeightWatch detects late-stage generalization collapse in grokking
The paper identifies anti-grokking, a late-stage collapse of generalization, in two extended grokking experiments: a 3-layer MLP on a subset of MNIST and a transformer trained on modular addition. The post says WeightWatcher detects signatures in layer weight matrices without training or test data.
It explores whether model weights can reveal a loss of generalization during prolonged training.
EB-JEPA: An Open-Source Library for Joint-Embedding Predictive Architectures
The paper presents EB-JEPA, an open-source library with modular implementations for learning representations and world models using JEPAs. These architectures predict in representation space rather than pixel space.
Engineers can explore implementations of JEPA-based representation learning for downstream tasks.
- Training and fine-tuningArticle
Book chapter introduces GRPO for verifiable-reward RL
A chapter of *Build a Reasoning Model (From Scratch)* covers GRPO from scratch for reinforcement learning with verifiable rewards. The author says follow-up material will explore additional runs, analyses, and algorithmic tweaks.
Useful for engineers seeking a step-by-step introduction to GRPO and reasoning-model training.
Build with AgentLog
Send your own newsletter
CommsHarbor keeps contacts, consent and one-click unsubscribe together.
Free workspace
Open CommsHarbor





