Skip to content
EN

Attention and model architecture

Transformers, attention variants, state-space models and the ideas behind new architectures.

156 links, newest first.

Get the weekly briefing

The best new links of the topics you pick, summarized with the source. At most one email a week.

Topics: Attention and model architecture

Before the first issue we email you to confirm; leaving takes one click. Sent with CommsHarbor. Privacy

  1. Grok’s Triton Paged Attention Kernel

    A post shares a paged attention kernel attributed to Grok, with a link to a solution file hosted on KernelBench.

    Engineers can inspect the implementation of a paged attention kernel.

  2. Three papers on efficient LLM serving

    The post recommends papers on PagedAttention, Sarathi-Serve, and SGLang, highlighting KV-cache management, chunked prefills, scheduling, and structured generation. The PagedAttention paper describes virtual-memory-inspired KV-cache management to reduce fragmentation and duplication.

    These papers cover techniques for managing KV-cache memory and improving LLM serving throughput and utilization.

  3. Building Diffusion Language Models

    A Kuleshov Group blog post introduces diffusion language models and research behind recent open-source models. It covers masking diffusion, iterative refinement, post-training, and variable-length generation.

    Engineers can use it to understand the techniques that make diffusion language models work.

  4. Qwen 3.5 and Nemotron 3 use different hybrid architectures

    The post describes Qwen 3.5 as a Gated Delta Net hybrid and linear-attention variant, and Nemotron 3 as a Mamba hybrid. It connects linear attention and state-space models as forms of convergent evolution.

    The comparison highlights architectural approaches beyond standard QKV attention.

  5. Mechanistic Data Attribution Traces Training Data Behind LLM Circuits

    The Mechanistic Data Attribution framework uses influence functions to link interpretable LLM units to training samples. The authors report that changing a few high-influence samples affects which heads emerge, and that interventions on induction heads alter in-context learning.

    It connects training data to the emergence and behavior of interpretable attention heads.

  6. HiLS learns chunk selection for sparse attention

    The paper proposes Hierarchical Landmark Sparse (HiLS) Attention, a chunk-wise sparse attention method that learns chunk selection end-to-end under the language-modeling loss. The post says chunks are scored using aggregated scores from landmark tokens.

    Engineers exploring long-context models can assess an approach to selecting chunks for sparse attention.

  7. HOLA combines recurrent state with a bounded KV cache

    The post describes HOLA, which pairs a recurrent state with a small KV cache for exact memory, routing tokens with high prediction error to the cache. It reports a 16.1% perplexity reduction and robust retrieval at 32k tokens.

    The design explores whether a bounded exact-memory buffer can improve long-context retrieval in linear-attention models.

  8. Functional Attention uses learned bases to reduce attention cost

    Functional Attention represents query and key/value fields with learned bases, then relates them through a compact matrix instead of an all-pairs score table. The paper reports linear cost in the number of points and results across PDE, point-cloud segmentation, and regression tasks.

    Engineers working on long sequences or operator learning can assess an alternative to quadratic all-pairs attention.

  9. HOLA Adds a 64-Token Cache to Linear Attention

    The post says HOLA uses a parameter-free token-surprise metric and a 64-token cache to address catastrophic forgetting in linear attention, claiming it beats full-attention baselines.

    The approach could matter to engineers balancing inference cost, memory recall, and cache size.

  10. FlowTracer Traces Attention-Based Information Flow for RL

    The paper introduces FlowTracer, which traces information propagation through LLM reasoning to identify influential tokens and uses those signals to target reinforcement-learning credit assignment.

    It offers a way to allocate RL training signals according to tokens that contribute to answer formation.

  11. HOLA Adds a Bounded Exact KV Cache to Linear Attention

    HOLA combines a delta-rule recurrent state with a bounded exact key-value cache. The paper reports improved Wikitext perplexity and robust RULER needle recall up to 32k tokens.

    The design offers a way to improve long-range recall while retaining linear attention’s compressed state.

  12. A Reading List on Transformers, Scaling, and Fine-Tuning

    The post lists ten papers and works, in a suggested order, covering transformers, scaling, few-shot learning, RLHF, LoRA, FlashAttention, chain-of-thought, and DPO.

    It offers engineers a structured path through influential LLM architecture and training concepts.

  13. A Study of Shared QKV Projections in Transformers

    The paper systematically evaluates Transformer variants that share key and value, query and key, or all three projections. The post reports that setting K equal to V was too much for the author's small ternary LLMs.

    Projection sharing can reduce the number of distinct attention projections, and the results may inform architecture choices.

  14. Video explains the transformer attention pipeline

    The video covers Q/K/V, attention scores, and encoder blocks. The post describes attention as matching each token’s query against sequence keys to weight value vectors.

    A concise walkthrough of the core operations behind transformer attention may help engineers build intuition for model architecture.

  15. Replacing Transformer Attention Heads with Synthesized Programs

    The post describes synthesizing Python programs from attention maps to reproduce individual heads’ token-to-token patterns. On GPT-2, TinyLlama, and Llama-3B, the best programs reached 69–79% mean IoU; replacing up to 30–40% of heads reportedly caused little QA decline.

    This explores whether some attention heads can be approximated by explicit, swappable programs rather than neural computations.

  16. HydraHead Mixes Full and Linear Attention by Head

    HydraHead is an attention hybridization architecture that combines Full Attention and Linear Attention at the head level. The post says it is motivated by mechanistic interpretability and aims to support more efficient long-context models.

    The head-level design offers an alternative granularity for combining attention mechanisms in long-context models.

  17. Explaining Transformer Attention Heads with Python Programs

    A paper introduces an automated interpretability technique that explains attention heads with Python programs. Replacing about 40% of attention patterns in Llama-3B with program outputs barely affects task performance.

    The technique could help engineers identify attention patterns that can be simplified or inform architectural changes.

  18. Explaining Transformer Attention Heads with Programs

    The paper proposes approximating deep-network components with executable programs, focusing on attention heads in transformer language models. It computes attention matrices on randomly selected training examples and prompts a pretrained language model.

    Programmatic approximations of attention heads could help engineers study how transformer components behave.

  19. How reasoning unlocks knowledge stored in LLMs

    A Google Research blog studies how reasoning helps LLMs recover knowledge stored in their weights.

    It examines how reasoning can affect access to an LLM’s parametric knowledge.

  20. Geometric Structures Behind Arithmetic in LLMs

    The paper analyzes residual-stream representations during multi-operand addition and identifies the Iso-Raw-Sum Trajectory, where semantic digits anchor representations and continuous carry fibers modulate them. It proposes the Noisy Quantization Model to explain arithmetic errors.

    It offers a geometric account of how LLM representations encode carries and where arithmetic errors may arise.

Build with AgentLog

List your MCP, skill or plugin

Reach the engineers who read these briefings.

Sponsor AgentLog

Footer, sidebar or featured slot for 30 days.

From US$ 60

See the slots

Send your own newsletter

CommsHarbor keeps contacts, consent and one-click unsubscribe together.

Free workspace

Open CommsHarbor