Training and fine-tuning
Pre-training, post-training, LoRA and the recipes behind better models.
49 links, newest first.
- Training and fine-tuningArticle
Self-Play Pretraining Without Natural Data
The paper trains a language model on byte sequences generated by programs, while an RL-updated generator adapts to the learner’s progress. It reports transfer to unseen text, images, audio, and code, along with in-context learning.
The approach explores whether self-generated data can teach transferable structure without natural-data gradient updates.
- Training and fine-tuningPost on X
Self-Play Pretraining with Zero Data
The work trains a generator and learner from random initialization: the generator proposes programs for a universal Turing machine, and the learner trains on their outputs. It reports predictable reductions in zero-shot validation loss across images, text, audio, and melodies, plus in-context…
It explores whether self-play can produce scalable pretraining without training on real data.
- Training and fine-tuningPost on X
A progression of policy-gradient methods for LLM training
The post outlines an evolution from vanilla policy gradient and REINFORCE to PPO, GRPO, and GRPO variants, describing how REINFORCE estimates the policy gradient using sampled rollouts.
It offers engineers a concise conceptual map of reinforcement-learning methods used to train LLMs.
LoRA+ uses different learning rates for adapter matrices
The paper analyzes how using the same learning rate for LoRA’s A and B matrices can hinder feature learning in wide models. It proposes using different learning rates to address this issue.
It offers a practical fine-tuning change for improving feature learning in wide models.
- Training and fine-tuningPost on X
Weight Spectra Track Random-Label Memorization in LLMs
A preliminary analysis uses WeightWatcher to monitor LLM weight spectra during training with random labels. The attention K matrix’s power-law exponent α tracks memorization across seeds, but the effect is layer-specific.
The findings suggest a possible spectral signal for tracking when training shifts from generalization toward memorization.
onPanda: Token-Level Annotation and Model Inspection
onPanda is an open-source tool for LLM data annotation and model inspection. Its workflow supports token-level corrections, SFT and preference data, and inspection of token probabilities and decoding.
Engineers working on alignment can explore a workflow for annotating on-policy data and debugging model outputs.
- Training and fine-tuningPost on X
TR-DPO Adds a KL Penalty to the DPO Loss
The post describes Trust Region DPO (TR-DPO), which adds a KL-divergence penalty directly to the DPO loss to limit deviation from the reference model. The author says this stabilizes offline RL training.
The approach may be relevant to engineers evaluating stability and distribution shift in DPO training.
QLoRA enables 65B model fine-tuning on a 48GB GPU
The QLoRA paper describes fine-tuning a 65B-parameter model on one 48GB GPU by backpropagating through a frozen 4-bit quantized model into LoRA adapters. It reports preserving full 16-bit fine-tuning task performance.
Engineers can use this approach to reduce GPU memory requirements for fine-tuning large language models.
- Training and fine-tuningPost on X
Using GEPA to Tune a Model with Prompts and Labeled Examples
The post describes a demo using GEPA to adapt the off-the-shelf Jev model, trained on synthetic data, to specific decision criteria by changing the prompt and labeling a few examples.
It outlines a prompt-and-labeling approach for adapting a model to specific decision criteria.
ORPO combines preference optimization with supervised fine-tuning
The paper introduces ORPO, a reference-model-free approach that combines preference alignment with supervised fine-tuning. It uses a penalty for disfavored generations during preference-aligned SFT.
Engineers can evaluate an alternative to separate SFT and preference-alignment stages that does not require a reference model.
- Training and fine-tuningRepository
GuppyLM: a 9M-parameter language model
The post points to GuppyLM, a roughly 9M-parameter language model described as talking like a small fish. The author says it can be trained from scratch in five minutes.
The repository may offer engineers a small-scale project to explore language-model training.
JEPA-Anything uses orthogonal factors for predictive modeling
JEPA-Anything is a domain-agnostic framework that uses orthogonal predictive factorization to decompose latent targets into complementary factors. The authors say code, checkpoints, and a report are open.
The paper describes a way to structure predictive-model capacity across different domains.
- Training and fine-tuningPost on X
TinyLoRA tests RL fine-tuning with 13 parameters
The post describes TinyLoRA, which replaces LoRA’s low-rank matrix with a trainable vector projected through a random tensor. It reports that GRPO training with 13 parameters improved Qwen2.5-7B-Instruct scores on GSM8K, MATH500, and AIME24.
It explores whether RL-based fine-tuning can reduce adapter parameter and memory requirements.
RoLA: Low-Rank Linear Attention for Diffusion Transformers
The paper introduces RoLA, a rotary-positioned low-rank linear attention method for Diffusion Transformers. It addresses the quadratic scaling of dense spatiotemporal self-attention and a compatibility issue between RoPE and the global branch of sparse low-rank hybrids.
Engineers working on video generation can assess an approach to reducing an inference bottleneck in Diffusion Transformers.
- Training and fine-tuningPost on X
NoRA changes LoRA initialization by normalizing A
The post describes Normalized Low-Rank Adaptation (NoRA), which normalizes columns of LoRA’s A matrix at initialization. It claims NoRA-init captures most of the gains without ongoing normalization and reports higher averages than several adapter methods on SFT and RLVR.
The initialization approach may offer a practical way to improve LoRA training without continuous normalization.
- Training and fine-tuningPost on X
Marin exposes details of its model training
The post says Marin provides views of its pre-training mixture by domain, sampled documents, live training loss, configs, and scaling laws.
These materials can help engineers inspect a model training run and its underlying choices.
- Training and fine-tuningPost on X
Ostris AI Toolkit Adds LoRA Training for MiniMax H3
Ostris AI Toolkit supports training LoRAs for MiniMax H3, currently limited to text-to-video and image-to-video. The post says it uses Comfy quantized weights.
Engineers working on video-model fine-tuning can check the toolkit’s current support and limitations.
- Training and fine-tuningPost on X
A First-Principles Guide to JEPA’s SIGReg
The post describes a guide to JEPA’s SIGReg that covers complex numbers, Euler’s formula, the Cramér–Wold theorem, and a working training loop.
It offers a step-by-step explanation of concepts and a training loop relevant to model training.
- Training and fine-tuningPost on X
The Stack v3 releases 5T deduplicated code tokens
The Stack v3 is an open code dataset with about 5T tokens of deduplicated and filtered source code across 713 languages. The post describes it as excluding restrictively licensed code.
Engineers training code models can evaluate a larger dataset with broad language coverage.
- Training and fine-tuningPost on X
Dual On-Policy Distillation Routes Supervision Per Token
The post describes a paper on Dual On-Policy Distillation, which dynamically routes each token between teacher-based and self-based supervision.
It may interest engineers exploring alternatives to standard teacher distillation and self-distillation.
Build with AgentLog
Send your own newsletter
CommsHarbor keeps contacts, consent and one-click unsubscribe together.
Free workspace
Open CommsHarbor

