Skip to content
EN

Briefings

Engineering links on agents, inference, retrieval and GPUs, picked by a person and summarized in two lines each, with the source.

Get the weekly briefing

The best new links of the topics you pick, summarized with the source. At most one email a week.

Topics: all

Before the first issue we email you to confirm; leaving takes one click. Sent with CommsHarbor. Privacy

Latest

  1. HySparse2 uses two-level KV sharing for sparse attention

    HySparse2 is a hybrid sparse attention method with two-level KV sharing. It targets efficient prefill and compact KV-cache storage for long-horizon, multi-turn agents.

    Engineers working on long-context models can assess an attention design aimed at reducing prefill and KV-cache demands.

  2. Memory Attention Reuses Keys in the Value Expression

    A post describes preliminary ablation results suggesting that reusing keys in V=K+M is an effective architectural choice. The author says it still needs validation at larger scales.

    The result may inform attention architecture experiments, but the post notes that larger-scale validation is still needed.

  3. LENS jointly trains reasoning and image segmentation with RL

    LENS is a reinforcement-learning framework for text-prompted image segmentation that jointly optimizes reasoning and segmentation. The paper says supervised fine-tuning often omits explicit chain-of-thought at test time, limiting generalization to unseen prompts and domains.

    Engineers working on vision systems can assess an approach that trains reasoning and segmentation together.

  4. Efficient Autoregressive Inference for Transformer Probabilistic Models

    An ICLR 2026 paper on efficient autoregressive inference for transformer probabilistic models, which the post says addresses a decoding bottleneck.

    Relevant to engineers working on faster autoregressive decoding for probabilistic transformer models.

  5. SecurityArticle

    Post claims iOS IPA decryption without a physical iPhone

    The post links to a report about decrypting FairPlay-protected iOS IPAs, claiming the process works without a physical iPhone or jailbreak.

    The claim may interest engineers studying iOS app protection and reverse engineering.

  6. PolaFormer++ Adds Polarity-Aware Features to Linear Attention

    PolaFormer++ introduces a theoretical criterion for feature-map spikiness and proposes a polarity-aware, channel-wise feature map. The authors evaluate linear attention across six visual task families.

    It offers a feature-map design and theoretical framework for engineers exploring linear attention in vision models.

  7. AI agentsArticle

    JAZ explores a minimalist recursive agent framework

    The post describes JAZ as using a single recursive code primitive for agent loops. It claims JAZ can outperform Letta and ACE at lower cost; the linked preview describes agents using tools and alternating through a workflow.

    The comparison may help engineers assess whether simpler agent loops can replace specialized memory and self-improvement frameworks.

  8. Dual-Cache enables cross-layer communication between LLMs

    The paper presents Dual-Cache, a latent protocol for transferring information between heterogeneous language models by translating a Sharer’s KV cache into a Receiver’s. The post says it adds joint cross-layer mixing.

    Engineers building multi-agent LLM systems can assess a way to share context without exchanging text.

  9. TWT Compresses Similar Vision Transformer Layers

    Transformer-Within-Transformer replaces contiguous groups of similar Vision Transformer layers with single attention-based surrogate layers. On DINOv2, the post reports roughly half the compute with a tiny accuracy drop.

    It describes a way to reduce Vision Transformer depth and compute while preserving most of the reported accuracy.

  10. JAZ uses a minimal, code-driven agent harness

    JAZ is an agent framework with a single `invoke` primitive; the LLM writes code that can call it recursively. The post reports results on StuLife and AppWorld, including lower costs than the compared approaches.

    The framework explores implementing memory and self-improvement within the agent loop instead of as separate subsystems.

  11. Self-Play Pretraining Without Natural Data

    The paper trains a language model on byte sequences generated by programs, while an RL-updated generator adapts to the learner’s progress. It reports transfer to unseen text, images, audio, and code, along with in-context learning.

    The approach explores whether self-generated data can teach transferable structure without natural-data gradient updates.

  12. Reducing Redundancy in Looped Transformers

    The paper identifies three types of computational redundancy across loops in looped Transformers. It reports 1.65× lower latency and 6× less memory by exploiting them.

    Engineers evaluating looped Transformers can assess techniques aimed at reducing their compute and memory costs.

  13. Ablations of Coding Agent Harnesses

    The paper reports 176 ablation setups across SWE-Bench and Terminal-Bench. The post says rule-based context trimming with a short summary outperformed elaborate recovery, while explicit planning helped weaker models but not frontier models’ success rates.

    Its ablations can help engineers assess which harness components improve agent performance or reduce costs.

  14. Unsloth guide to reinforcement learning with GRPO

    Unsloth documentation presents a beginner-to-advanced reinforcement learning guide, including how to train a DeepSeek-R1 reasoning model with GRPO.

    Engineers can use it as a guide to applying GRPO to reasoning-model training with Unsloth.

  15. Native Sparse Attention for Efficient Long-Context Modeling

    The paper presents NSA, a natively trainable sparse attention mechanism that combines algorithmic innovations with hardware-aligned optimizations for efficient long-context modeling.

    It describes an approach to sparse attention designed to improve efficiency while maintaining model capabilities.

  16. SVD-LLM Uses Truncation-Aware SVD for LLM Compression

    The paper proposes an SVD-based method for LLM compression that addresses compression loss from truncating smaller singular values and updates compressed weights after truncation.

    It describes an alternative to quantization for reducing model size, relevant to practical LLM deployment.

  17. LLM inferencePost on X

    How vLLM Schedules Requests and Allocates KV Cache

    An article traces a vLLM request from queueing through prefill and decoding. It explains how iteration-level scheduling and PagedAttention allocate KV cache blocks as sequences grow.

    Useful for understanding how vLLM manages batching and KV cache memory during inference.

  18. Xiaomi MiMo releases RL environments and training code

    A post points to Xiaomi MiMo’s GitHub repository for RL environments and training code, plus a MiMo-V2.6 RL dataset on Hugging Face.

    The repository and dataset may help engineers explore or build RL workflows for language models.

  19. Sparse Layers Are Critical to Scaling Looped Language Models

    The paper compares standard and Mixture-of-Experts transformers, with and without looping. It reports that Looped-MoE models scale better than the standard baseline, while dense looped models do not.

    The results help engineers assess sparse layers and looping when designing models with adaptive depth.

  20. Self-Play Pretraining with Zero Data

    The work trains a generator and learner from random initialization: the generator proposes programs for a universal Turing machine, and the learner trains on their outputs. It reports predictable reductions in zero-shot validation loss across images, text, audio, and melodies, plus in-context…

    It explores whether self-play can produce scalable pretraining without training on real data.

  21. SmolDataEnvs: 5,000 verifiable RL tasks

    SmolDataEnvs is an open-source collection of 5,000 verifiable reinforcement-learning environment tasks for hill-climbing small models in code and data science. The release includes environments, evaluations, and training.

    Engineers can use the tasks and accompanying evaluations to train and assess small models.

  22. Memory Attention Adds Token-Indexed Memory to Attention

    Memory Attention replaces the Transformer’s learned value projection with a sum of contextual keys and layer-specific, token-indexed memory vectors. The post says the memory can reside on CPU.

    Engineers exploring attention architectures can assess an approach that adds model capacity through token-indexed memory.

  23. Memory Attention replaces the value projection with token memory

    Memory Attention combines layer-specific, token-indexed memory with contextual keys to form attention values. The post says this makes value computation largely lookup-based, with potential CPU offloading and KV-cache reduction.

    Engineers evaluating attention alternatives can assess its memory, compute, and KV-cache trade-offs.

  24. LLM inferenceRepository

    NVIDIA Model Optimizer combines model compression techniques

    NVIDIA Model Optimizer is a library for techniques including quantization, distillation, pruning, neural architecture search, and speculative decoding. The post says its exported checkpoints can be used with vLLM, SGLang, TensorRT-LLM, and TensorRT.

    Engineers can evaluate a single toolkit for reducing model size and preparing checkpoints for inference runtimes.

  25. A progression of policy-gradient methods for LLM training

    The post outlines an evolution from vanilla policy gradient and REINFORCE to PPO, GRPO, and GRPO variants, describing how REINFORCE estimates the policy gradient using sampled rollouts.

    It offers engineers a concise conceptual map of reinforcement-learning methods used to train LLMs.

  26. RetNet combines parallel training with recurrent inference

    The Retentive Network paper proposes an architecture for language models with a retention mechanism and parallel, recurrent, and chunkwise recurrent computation paradigms. It derives a connection between recurrence and attention.

    Its alternative sequence-modeling mechanism and inference paradigms are relevant to engineers evaluating Transformer architectures.

  27. PagedAttention manages KV cache memory for LLM serving

    The paper introduces PagedAttention, an attention algorithm inspired by virtual memory, to address KV cache fragmentation and redundant duplication that can limit batch size in LLM serving.

    It describes a memory-management approach that can help engineers understand constraints on serving batch size.

  28. GraphRAG for query-focused summarization

    The paper presents GraphRAG, an approach for answering questions over document collections. It addresses global questions about a corpus, which standard RAG retrieval does not handle well.

    It describes an approach for corpus-level questions that may not be answered by retrieving individual passages.

  29. FlashAttention-3 targets faster attention on Hopper GPUs

    The paper presents techniques to speed up attention on Hopper GPUs, including exploiting asynchrony and low-precision computation. It addresses GPU memory traffic and hardware utilization.

    Engineers working on Transformer performance can assess attention optimizations designed for Hopper GPUs.

  30. How Mamba and Transformers Connect Through State-Space Duality

    The paper develops theoretical connections between state-space models such as Mamba and variants of attention, using decompositions of structured semiseparable matrices.

    The framework helps engineers understand the relationship between SSMs and attention architectures.

Build with AgentLog

List your MCP, skill or plugin

Reach the engineers who read these briefings.

Sponsor AgentLog

Footer, sidebar or featured slot for 30 days.

From US$ 60

See the slots

Send your own newsletter

CommsHarbor keeps contacts, consent and one-click unsubscribe together.

Free workspace

Open CommsHarbor