Training and fine-tuning
Pre-training, post-training, LoRA and the recipes behind better models.
49 links, newest first.
- Training and fine-tuningPost on X
ACE Accumulates Structured Context Using Performance Feedback
The post describes Agentic Context Engineering (ACE), which incrementally accumulates and refines structured context items using performance feedback. It contrasts this approach with GEPA's iterative prompt rewriting.
The approach offers an alternative to rewriting prompts when optimizing context for complex tasks.
- Training and fine-tuningPost on X
JEPA-SCORE Uses JEPA Encoders for Density Estimation
A study reports that the anti-collapse term in JEPAs implicitly estimates data density. It introduces JEPA-SCORE, which uses an encoder’s Jacobian to estimate sample probabilities without retraining.
It suggests a way to reuse self-supervised encoders for data curation and outlier detection.
- Training and fine-tuningPost on X
UserRL trains user-centric agents with simulated users
Salesforce proposes UserRL, a framework using standardized gym environments and simulated users to train agentic models. The post reports findings on SFT cold starts, trajectory-level rewards, and scalable simulated-user training.
The reported training choices may inform designs for multi-turn agent training.
- Training and fine-tuningPost on X
GEPA Prompts for Detecting Malicious AI-Generated Code
A post reports that GEPA found prompts that detected malicious behavior in AI-generated code, blocking 90% of malicious submissions at 1% of the audit budget.
It describes a prompt-optimization use case for evaluating AI-generated code.
- Training and fine-tuningPost on X
LoRA recipes for reinforcement learning fine-tuning
A Thinky Machines post describes using 10× larger learning rates and applying LoRA to all layers; it claims rank-1 LoRA can match full-finetuning performance when configured well.
The post highlights LoRA configuration choices engineers can evaluate for RL fine-tuning.
- Training and fine-tuningPost on X
SWE-Swiss training recipe for software engineering models
Researchers from Peking University, ByteDance Seed, and HKU introduce SWE-Swiss, a training recipe for software engineering tasks. They report that the 32B-parameter SWE-Swiss-32B scores 60.2% on SWE-bench Verified.
The reported results may help engineers assess training approaches for software engineering models.
- Training and fine-tuningPost on X
PEER uses product keys to retrieve from over a million experts
PEER is a Mixture of Experts architecture with product-key routing and single-neuron MLP experts. It retrieves a small set of experts for each input and combines their outputs with router scores.
The design offers engineers an approach to scaling model capacity through sparse expert retrieval.
- Training and fine-tuningPost on X
A PyTorch LoRA Fine-Tuning Example for a 5M-Parameter MLP
The author implemented LoRA to fine-tune a simple MLP for classification. The exercise freezes the original weights, trains about 9,000 adapter parameters on one label class, and compares performance with the adapter enabled and disabled.
It gives engineers a compact example of applying LoRA adapters and evaluating their effect.
- Training and fine-tuningPost on X
Fine-tuning Llama 3 8B with knowledge generated by Llama 3 405B
An open-source Colab notebook fine-tunes Llama 3 8B on domain-specific knowledge generated by Llama 3 405B.
Engineers can use the notebook as a starting point for distilling domain knowledge from a larger model into a smaller one.
Build with AgentLog
Send your own newsletter
CommsHarbor keeps contacts, consent and one-click unsubscribe together.
Free workspace
Open CommsHarbor