Skip to content
EN

Training and fine-tuning

Pre-training, post-training, LoRA and the recipes behind better models.

108 links, newest first.

Get the weekly briefing

The best new links of the topics you pick, summarized with the source. At most one email a week.

Topics: Training and fine-tuning

Before the first issue we email you to confirm; leaving takes one click. Sent with CommsHarbor. Privacy

  1. A GRPO stability setting from the INTELLECT-2 report

    The post says the INTELLECT-2 tech report introduced a setting to increase GRPO training stability: pass `delta=4.0`.

    It points to a concrete training parameter engineers can investigate when tuning GRPO.

  2. Unsloth guide to fine-tuning embedding models

    Unsloth documentation explains how to fine-tune embedding models. The post links to a Colab notebook for EmbeddingGemma (300M).

    Useful for engineers adapting embedding models for retrieval and semantic search.

  3. DroPE removes positional embeddings for context extension

    The post describes DroPE, which drops positional embeddings after training and uses short calibration. It claims zero-shot context extension and improved 16k retrieval and Q&A over baselines for models up to 7B.

    The approach may offer a way to extend context without long-context fine-tuning.

  4. South Korea’s Model Training Grant Sets Three Constraints

    The first round has five teams, with four advancing. Projects must train from scratch, be somewhat open for commercial use, and pursue ambitious goals such as large-scale omni models.

    The constraints offer a snapshot of how a public program is structuring model-training competition.

  5. VL-JEPA Predicts Text Embeddings for Vision-Language Tasks

    VL-JEPA is a vision-language model based on JEPA that predicts continuous embeddings of target text rather than generating tokens autoregressively. The paper says this approach emphasizes task-relevant semantics over surface-level wording variation.

    The paper explores an alternative training objective for vision-language models that may reduce sensitivity to wording variation.

  6. Character.AI’s Squinch Gradient Compression for Pretraining

    The post says Character.AI used Squinch, a gradient compression algorithm, to maintain SOTA MFU while pretraining on GCP H100-TCPX with one-quarter the bandwidth of IB.

    It highlights how gradient compression can help sustain training efficiency under network bandwidth constraints.

  7. LeJEPA combines JEPA prediction loss with SIGReg

    The post describes LeJEPA as combining JEPA’s latent-space prediction loss with SIGReg, a regularizer intended to shape embeddings into an isotropic Gaussian distribution. The linked article discusses its theoretical foundations and practical use.

    Useful for engineers exploring self-supervised learning recipes for JEPA models.

  8. RobustMerge combines fine-tuned MLLMs without retraining

    RobustMerge is a training-free, parameter-efficient method for merging fine-tuned expert models into a multi-task MLLM. The post says it uses low-rank analysis and cross-task normalization to preserve direction robustness.

    It offers an approach to consolidate task-specific models without retraining them.

  9. CE-GPPO preserves clipped-token gradients in LLM reinforcement learning

    The paper introduces CE-GPPO, an RL method for LLM training that addresses gradient signals from low-probability tokens discarded by PPO-style clipping. It focuses on coordinating policy entropy and exploration–exploitation.

    It offers engineers a PPO alternative to investigate for managing entropy during LLM reinforcement learning.

  10. Ant Group introduces Ring-linear 2.0 hybrid models

    The post describes Ring-mini-linear-2.0 (16B) and Ring-flash-linear-2.0 (104B), which combine linear and softmax attention. It claims lower inference costs and 50% better training efficiency using a custom FP8 operator library.

    The models’ attention design and FP8 training approach may be relevant to engineers optimizing long-context inference and training.

  11. SFT Can Preserve General Capabilities with Smaller Learning Rates

    The study revisits domain-specific fine-tuning in LLMs and reports that smaller learning rates can preserve general skills while maintaining domain performance. It introduces Token-Adaptive Loss Reweighting (TALR), which the post says outperforms LoRA, FLOW, and other baselines.

    Engineers tuning models for specialized tasks can use the findings to balance domain performance and general capabilities.

  12. Text-to-LoRA Generates Adapters From Natural-Language Descriptions

    Sakana AI’s Text-to-LoRA generates task-specific LoRA adapters for Mistral-7B-Instruct from text descriptions. Trained on 479 tasks, it averaged 67.7% accuracy, outperforming the base model but slightly trailing conventional task-specific adapters.

    It offers an approach to producing task-specific adapters on demand without training a new adapter for each task.

  13. Hugging Face Course on Prompt Optimization with DSPy GEPA

    Hugging Face offers a course on prompt optimization for language models using DSPy GEPA.

    It gives engineers a learning resource on optimizing prompts for language models.

  14. Prompt Optimization with DSPy GEPA

    A Hugging Face cookbook on using DSPy GEPA for reflective prompt optimization. The post reports an 11% accuracy boost on NuminaMath-1.5 for under $0.50 total cost.

    It offers an example of optimizing prompts with separate models for inference and reflection.

  15. ACE Accumulates Structured Context Using Performance Feedback

    The post describes Agentic Context Engineering (ACE), which incrementally accumulates and refines structured context items using performance feedback. It contrasts this approach with GEPA's iterative prompt rewriting.

    The approach offers an alternative to rewriting prompts when optimizing context for complex tasks.

  16. A guide to LLM post-training

    The guide covers the transition from next-token prediction to instruction following, including SFT data and objectives, RLHF, RLAIF, RLVR, reward models, and evaluation frameworks.

    It outlines post-training methods and evaluation topics engineers may need when adapting LLMs.

  17. JEPA-SCORE Uses JEPA Encoders for Density Estimation

    A study reports that the anti-collapse term in JEPAs implicitly estimates data density. It introduces JEPA-SCORE, which uses an encoder’s Jacobian to estimate sample probabilities without retraining.

    It suggests a way to reuse self-supervised encoders for data curation and outlier detection.

  18. UserRL trains user-centric agents with simulated users

    Salesforce proposes UserRL, a framework using standardized gym environments and simulated users to train agentic models. The post reports findings on SFT cold starts, trajectory-level rewards, and scalable simulated-user training.

    The reported training choices may inform designs for multi-turn agent training.

  19. GEPA Prompts for Detecting Malicious AI-Generated Code

    A post reports that GEPA found prompts that detected malicious behavior in AI-generated code, blocking 90% of malicious submissions at 1% of the audit budget.

    It describes a prompt-optimization use case for evaluating AI-generated code.

  20. LoRA recipes for reinforcement learning fine-tuning

    A Thinky Machines post describes using 10× larger learning rates and applying LoRA to all layers; it claims rank-1 LoRA can match full-finetuning performance when configured well.

    The post highlights LoRA configuration choices engineers can evaluate for RL fine-tuning.

Build with AgentLog

List your MCP, skill or plugin

Reach the engineers who read these briefings.

Sponsor AgentLog

Footer, sidebar or featured slot for 30 days.

From US$ 60

See the slots

Send your own newsletter

CommsHarbor keeps contacts, consent and one-click unsubscribe together.

Free workspace

Open CommsHarbor