Briefings
Engineering links on agents, inference, retrieval and GPUs, picked by a person and summarized in two lines each, with the source.
Latest
HySparse2 uses two-level KV sharing for sparse attention
HySparse2 is a hybrid sparse attention method with two-level KV sharing. It targets efficient prefill and compact KV-cache storage for long-horizon, multi-turn agents.
Engineers working on long-context models can assess an attention design aimed at reducing prefill and KV-cache demands.
- Attention and model architecturePost on X
Memory Attention Reuses Keys in the Value Expression
A post describes preliminary ablation results suggesting that reusing keys in V=K+M is an effective architectural choice. The author says it still needs validation at larger scales.
The result may inform attention architecture experiments, but the post notes that larger-scale validation is still needed.
LENS jointly trains reasoning and image segmentation with RL
LENS is a reinforcement-learning framework for text-prompted image segmentation that jointly optimizes reasoning and segmentation. The paper says supervised fine-tuning often omits explicit chain-of-thought at test time, limiting generalization to unseen prompts and domains.
Engineers working on vision systems can assess an approach that trains reasoning and segmentation together.
Efficient Autoregressive Inference for Transformer Probabilistic Models
An ICLR 2026 paper on efficient autoregressive inference for transformer probabilistic models, which the post says addresses a decoding bottleneck.
Relevant to engineers working on faster autoregressive decoding for probabilistic transformer models.
- SecurityArticle
Post claims iOS IPA decryption without a physical iPhone
The post links to a report about decrypting FairPlay-protected iOS IPAs, claiming the process works without a physical iPhone or jailbreak.
The claim may interest engineers studying iOS app protection and reverse engineering.
PolaFormer++ Adds Polarity-Aware Features to Linear Attention
PolaFormer++ introduces a theoretical criterion for feature-map spikiness and proposes a polarity-aware, channel-wise feature map. The authors evaluate linear attention across six visual task families.
It offers a feature-map design and theoretical framework for engineers exploring linear attention in vision models.
- AI agentsArticle
JAZ explores a minimalist recursive agent framework
The post describes JAZ as using a single recursive code primitive for agent loops. It claims JAZ can outperform Letta and ACE at lower cost; the linked preview describes agents using tools and alternating through a workflow.
The comparison may help engineers assess whether simpler agent loops can replace specialized memory and self-improvement frameworks.
- LLM inferencePaper
Dual-Cache enables cross-layer communication between LLMs
The paper presents Dual-Cache, a latent protocol for transferring information between heterogeneous language models by translating a Sharer’s KV cache into a Receiver’s. The post says it adds joint cross-layer mixing.
Engineers building multi-agent LLM systems can assess a way to share context without exchanging text.
TWT Compresses Similar Vision Transformer Layers
Transformer-Within-Transformer replaces contiguous groups of similar Vision Transformer layers with single attention-based surrogate layers. On DINOv2, the post reports roughly half the compute with a tiny accuracy drop.
It describes a way to reduce Vision Transformer depth and compute while preserving most of the reported accuracy.
- AI agentsPaper
JAZ uses a minimal, code-driven agent harness
JAZ is an agent framework with a single `invoke` primitive; the LLM writes code that can call it recursively. The post reports results on StuLife and AppWorld, including lower costs than the compared approaches.
The framework explores implementing memory and self-improvement within the agent loop instead of as separate subsystems.
- Training and fine-tuningArticle
Self-Play Pretraining Without Natural Data
The paper trains a language model on byte sequences generated by programs, while an RL-updated generator adapts to the learner’s progress. It reports transfer to unseen text, images, audio, and code, along with in-context learning.
The approach explores whether self-generated data can teach transferable structure without natural-data gradient updates.
Reducing Redundancy in Looped Transformers
The paper identifies three types of computational redundancy across loops in looped Transformers. It reports 1.65× lower latency and 6× less memory by exploiting them.
Engineers evaluating looped Transformers can assess techniques aimed at reducing their compute and memory costs.
- AI agentsPaper
Ablations of Coding Agent Harnesses
The paper reports 176 ablation setups across SWE-Bench and Terminal-Bench. The post says rule-based context trimming with a short summary outperformed elaborate recovery, while explicit planning helped weaker models but not frontier models’ success rates.
Its ablations can help engineers assess which harness components improve agent performance or reduce costs.
- Reinforcement learningArticle
Unsloth guide to reinforcement learning with GRPO
Unsloth documentation presents a beginner-to-advanced reinforcement learning guide, including how to train a DeepSeek-R1 reasoning model with GRPO.
Engineers can use it as a guide to applying GRPO to reasoning-model training with Unsloth.
Native Sparse Attention for Efficient Long-Context Modeling
The paper presents NSA, a natively trainable sparse attention mechanism that combines algorithmic innovations with hardware-aligned optimizations for efficient long-context modeling.
It describes an approach to sparse attention designed to improve efficiency while maintaining model capabilities.
- LLM inferencePaper
SVD-LLM Uses Truncation-Aware SVD for LLM Compression
The paper proposes an SVD-based method for LLM compression that addresses compression loss from truncating smaller singular values and updates compressed weights after truncation.
It describes an alternative to quantization for reducing model size, relevant to practical LLM deployment.
- LLM inferencePost on X
How vLLM Schedules Requests and Allocates KV Cache
An article traces a vLLM request from queueing through prefill and decoding. It explains how iteration-level scheduling and PagedAttention allocate KV cache blocks as sequences grow.
Useful for understanding how vLLM manages batching and KV cache memory during inference.
- Reinforcement learningRepository
Xiaomi MiMo releases RL environments and training code
A post points to Xiaomi MiMo’s GitHub repository for RL environments and training code, plus a MiMo-V2.6 RL dataset on Hugging Face.
The repository and dataset may help engineers explore or build RL workflows for language models.
Sparse Layers Are Critical to Scaling Looped Language Models
The paper compares standard and Mixture-of-Experts transformers, with and without looping. It reports that Looped-MoE models scale better than the standard baseline, while dense looped models do not.
The results help engineers assess sparse layers and looping when designing models with adaptive depth.
- Training and fine-tuningPost on X
Self-Play Pretraining with Zero Data
The work trains a generator and learner from random initialization: the generator proposes programs for a universal Turing machine, and the learner trains on their outputs. It reports predictable reductions in zero-shot validation loss across images, text, audio, and melodies, plus in-context…
It explores whether self-play can produce scalable pretraining without training on real data.
- Reinforcement learningDataset
SmolDataEnvs: 5,000 verifiable RL tasks
SmolDataEnvs is an open-source collection of 5,000 verifiable reinforcement-learning environment tasks for hill-climbing small models in code and data science. The release includes environments, evaluations, and training.
Engineers can use the tasks and accompanying evaluations to train and assess small models.
Memory Attention Adds Token-Indexed Memory to Attention
Memory Attention replaces the Transformer’s learned value projection with a sum of contextual keys and layer-specific, token-indexed memory vectors. The post says the memory can reside on CPU.
Engineers exploring attention architectures can assess an approach that adds model capacity through token-indexed memory.
Memory Attention replaces the value projection with token memory
Memory Attention combines layer-specific, token-indexed memory with contextual keys to form attention values. The post says this makes value computation largely lookup-based, with potential CPU offloading and KV-cache reduction.
Engineers evaluating attention alternatives can assess its memory, compute, and KV-cache trade-offs.
- LLM inferenceRepository
NVIDIA Model Optimizer combines model compression techniques
NVIDIA Model Optimizer is a library for techniques including quantization, distillation, pruning, neural architecture search, and speculative decoding. The post says its exported checkpoints can be used with vLLM, SGLang, TensorRT-LLM, and TensorRT.
Engineers can evaluate a single toolkit for reducing model size and preparing checkpoints for inference runtimes.
- Training and fine-tuningPost on X
A progression of policy-gradient methods for LLM training
The post outlines an evolution from vanilla policy gradient and REINFORCE to PPO, GRPO, and GRPO variants, describing how REINFORCE estimates the policy gradient using sampled rollouts.
It offers engineers a concise conceptual map of reinforcement-learning methods used to train LLMs.
RetNet combines parallel training with recurrent inference
The Retentive Network paper proposes an architecture for language models with a retention mechanism and parallel, recurrent, and chunkwise recurrent computation paradigms. It derives a connection between recurrence and attention.
Its alternative sequence-modeling mechanism and inference paradigms are relevant to engineers evaluating Transformer architectures.
- LLM inferencePaper
PagedAttention manages KV cache memory for LLM serving
The paper introduces PagedAttention, an attention algorithm inspired by virtual memory, to address KV cache fragmentation and redundant duplication that can limit batch size in LLM serving.
It describes a memory-management approach that can help engineers understand constraints on serving batch size.
- RAG and retrievalPaper
GraphRAG for query-focused summarization
The paper presents GraphRAG, an approach for answering questions over document collections. It addresses global questions about a corpus, which standard RAG retrieval does not handle well.
It describes an approach for corpus-level questions that may not be answered by retrieving individual passages.
FlashAttention-3 targets faster attention on Hopper GPUs
The paper presents techniques to speed up attention on Hopper GPUs, including exploiting asynchrony and low-precision computation. It addresses GPU memory traffic and hardware utilization.
Engineers working on Transformer performance can assess attention optimizations designed for Hopper GPUs.
How Mamba and Transformers Connect Through State-Space Duality
The paper develops theoretical connections between state-space models such as Mamba and variants of attention, using decompositions of structured semiseparable matrices.
The framework helps engineers understand the relationship between SSMs and attention architectures.
Build with AgentLog
Send your own newsletter
CommsHarbor keeps contacts, consent and one-click unsubscribe together.
Free workspace
Open CommsHarbor




