Skip to content
EN

Training and fine-tuning

Pre-training, post-training, LoRA and the recipes behind better models.

49 links, newest first.

Get the weekly briefing

The best new links of the topics you pick, summarized with the source. At most one email a week.

Topics: Training and fine-tuning

Before the first issue we email you to confirm; leaving takes one click. Sent with CommsHarbor. Privacy

  1. ACE Accumulates Structured Context Using Performance Feedback

    The post describes Agentic Context Engineering (ACE), which incrementally accumulates and refines structured context items using performance feedback. It contrasts this approach with GEPA's iterative prompt rewriting.

    The approach offers an alternative to rewriting prompts when optimizing context for complex tasks.

  2. JEPA-SCORE Uses JEPA Encoders for Density Estimation

    A study reports that the anti-collapse term in JEPAs implicitly estimates data density. It introduces JEPA-SCORE, which uses an encoder’s Jacobian to estimate sample probabilities without retraining.

    It suggests a way to reuse self-supervised encoders for data curation and outlier detection.

  3. UserRL trains user-centric agents with simulated users

    Salesforce proposes UserRL, a framework using standardized gym environments and simulated users to train agentic models. The post reports findings on SFT cold starts, trajectory-level rewards, and scalable simulated-user training.

    The reported training choices may inform designs for multi-turn agent training.

  4. GEPA Prompts for Detecting Malicious AI-Generated Code

    A post reports that GEPA found prompts that detected malicious behavior in AI-generated code, blocking 90% of malicious submissions at 1% of the audit budget.

    It describes a prompt-optimization use case for evaluating AI-generated code.

  5. LoRA recipes for reinforcement learning fine-tuning

    A Thinky Machines post describes using 10× larger learning rates and applying LoRA to all layers; it claims rank-1 LoRA can match full-finetuning performance when configured well.

    The post highlights LoRA configuration choices engineers can evaluate for RL fine-tuning.

  6. SWE-Swiss training recipe for software engineering models

    Researchers from Peking University, ByteDance Seed, and HKU introduce SWE-Swiss, a training recipe for software engineering tasks. They report that the 32B-parameter SWE-Swiss-32B scores 60.2% on SWE-bench Verified.

    The reported results may help engineers assess training approaches for software engineering models.

  7. PEER uses product keys to retrieve from over a million experts

    PEER is a Mixture of Experts architecture with product-key routing and single-neuron MLP experts. It retrieves a small set of experts for each input and combines their outputs with router scores.

    The design offers engineers an approach to scaling model capacity through sparse expert retrieval.

  8. A PyTorch LoRA Fine-Tuning Example for a 5M-Parameter MLP

    The author implemented LoRA to fine-tune a simple MLP for classification. The exercise freezes the original weights, trains about 9,000 adapter parameters on one label class, and compares performance with the adapter enabled and disabled.

    It gives engineers a compact example of applying LoRA adapters and evaluating their effect.

  9. Fine-tuning Llama 3 8B with knowledge generated by Llama 3 405B

    An open-source Colab notebook fine-tunes Llama 3 8B on domain-specific knowledge generated by Llama 3 405B.

    Engineers can use the notebook as a starting point for distilling domain knowledge from a larger model into a smaller one.

Build with AgentLog

List your MCP, skill or plugin

Reach the engineers who read these briefings.

Sponsor AgentLog

Footer, sidebar or featured slot for 30 days.

From US$ 60

See the slots

Send your own newsletter

CommsHarbor keeps contacts, consent and one-click unsubscribe together.

Free workspace

Open CommsHarbor