Attention and model architecture
Transformers, attention variants, state-space models and the ideas behind new architectures.
156 links, newest first.
HySparse2 uses two-level KV sharing for sparse attention
HySparse2 is a hybrid sparse attention method with two-level KV sharing. It targets efficient prefill and compact KV-cache storage for long-horizon, multi-turn agents.
Engineers working on long-context models can assess an attention design aimed at reducing prefill and KV-cache demands.
- Attention and model architecturePost on X
Memory Attention Reuses Keys in the Value Expression
A post describes preliminary ablation results suggesting that reusing keys in V=K+M is an effective architectural choice. The author says it still needs validation at larger scales.
The result may inform attention architecture experiments, but the post notes that larger-scale validation is still needed.
Efficient Autoregressive Inference for Transformer Probabilistic Models
An ICLR 2026 paper on efficient autoregressive inference for transformer probabilistic models, which the post says addresses a decoding bottleneck.
Relevant to engineers working on faster autoregressive decoding for probabilistic transformer models.
PolaFormer++ Adds Polarity-Aware Features to Linear Attention
PolaFormer++ introduces a theoretical criterion for feature-map spikiness and proposes a polarity-aware, channel-wise feature map. The authors evaluate linear attention across six visual task families.
It offers a feature-map design and theoretical framework for engineers exploring linear attention in vision models.
TWT Compresses Similar Vision Transformer Layers
Transformer-Within-Transformer replaces contiguous groups of similar Vision Transformer layers with single attention-based surrogate layers. On DINOv2, the post reports roughly half the compute with a tiny accuracy drop.
It describes a way to reduce Vision Transformer depth and compute while preserving most of the reported accuracy.
Reducing Redundancy in Looped Transformers
The paper identifies three types of computational redundancy across loops in looped Transformers. It reports 1.65× lower latency and 6× less memory by exploiting them.
Engineers evaluating looped Transformers can assess techniques aimed at reducing their compute and memory costs.
Native Sparse Attention for Efficient Long-Context Modeling
The paper presents NSA, a natively trainable sparse attention mechanism that combines algorithmic innovations with hardware-aligned optimizations for efficient long-context modeling.
It describes an approach to sparse attention designed to improve efficiency while maintaining model capabilities.
Sparse Layers Are Critical to Scaling Looped Language Models
The paper compares standard and Mixture-of-Experts transformers, with and without looping. It reports that Looped-MoE models scale better than the standard baseline, while dense looped models do not.
The results help engineers assess sparse layers and looping when designing models with adaptive depth.
Memory Attention Adds Token-Indexed Memory to Attention
Memory Attention replaces the Transformer’s learned value projection with a sum of contextual keys and layer-specific, token-indexed memory vectors. The post says the memory can reside on CPU.
Engineers exploring attention architectures can assess an approach that adds model capacity through token-indexed memory.
Memory Attention replaces the value projection with token memory
Memory Attention combines layer-specific, token-indexed memory with contextual keys to form attention values. The post says this makes value computation largely lookup-based, with potential CPU offloading and KV-cache reduction.
Engineers evaluating attention alternatives can assess its memory, compute, and KV-cache trade-offs.
RetNet combines parallel training with recurrent inference
The Retentive Network paper proposes an architecture for language models with a retention mechanism and parallel, recurrent, and chunkwise recurrent computation paradigms. It derives a connection between recurrence and attention.
Its alternative sequence-modeling mechanism and inference paradigms are relevant to engineers evaluating Transformer architectures.
FlashAttention-3 targets faster attention on Hopper GPUs
The paper presents techniques to speed up attention on Hopper GPUs, including exploiting asynchrony and low-precision computation. It addresses GPU memory traffic and hardware utilization.
Engineers working on Transformer performance can assess attention optimizations designed for Hopper GPUs.
How Mamba and Transformers Connect Through State-Space Duality
The paper develops theoretical connections between state-space models such as Mamba and variants of attention, using decompositions of structured semiseparable matrices.
The framework helps engineers understand the relationship between SSMs and attention architectures.
Megalodon: Efficient Sequence Modeling with Unlimited Context
The paper introduces Megalodon, a sequence-modeling architecture designed for efficient pretraining and inference with unlimited context length. It builds on Mega and targets the long-sequence limitations of Transformers.
It offers engineers an architecture to evaluate for long-context workloads and efficient sequence modeling.
Differential Transformer subtracts two attention maps
The Differential Transformer subtracts two attention maps to reduce attention noise and focus on signal, according to the post.
The attention variant may be relevant to engineers exploring alternatives to standard Transformer attention.
Memory Attention Replaces Learned Value Projections
Memory Attention (MA) replaces the Transformer’s learned value projection with a sum of contextual keys and layer-specific, token-indexed memory vectors. The paper examines MA across several attention variants.
The design explores an alternative way to represent and retrieve information in attention layers.
HySparse2 uses KV sharing for long-context inference
HySparse2 introduces two levels of KV sharing: cross-decoder KV bridging and reuse between sparse and full-attention layers. The post reports lower prefill FLOPs, a smaller KV cache, and better retrieval scores than MiMo-V2.6's Hybrid SWA architecture at 1M tokens.
Its KV-sharing design targets prefill cost, cache size, and retrieval accuracy in growing-context agentic workloads.
- Attention and model architectureRepository
iSDFT: Information-Proximal Self-Distillation for Continual Learning
iSDFT is a method for continual learning in LLMs, described as information-proximal on-policy self-distillation. Its GitHub repository provides training code, checkpoints, and evaluation results.
The implementation and evaluation results offer engineers resources for exploring continual learning in LLMs.
Mixture-of-Depths Dynamically Allocates Transformer Compute
The paper presents a method for transformer language models to allocate compute to specific sequence positions across layers. It caps how many tokens participate in self-attention and MLP computations at each layer.
Engineers can explore an approach to varying per-token compute under a layer-level budget.
Quiet-STaR trains language models to think before speaking
Quiet-STaR explores teaching language models to generate internal reasoning before producing text, building on the idea that reasoning is implicit in much written language.
It offers engineers an approach to training models to reason before generating their visible output.
Build with AgentLog
Send your own newsletter
CommsHarbor keeps contacts, consent and one-click unsubscribe together.
Free workspace
Open CommsHarbor
