Attention and model architecture
Transformers, attention variants, state-space models and the ideas behind new architectures.
156 links, newest first.
Grok’s Triton Paged Attention Kernel
A post shares a paged attention kernel attributed to Grok, with a link to a solution file hosted on KernelBench.
Engineers can inspect the implementation of a paged attention kernel.
Three papers on efficient LLM serving
The post recommends papers on PagedAttention, Sarathi-Serve, and SGLang, highlighting KV-cache management, chunked prefills, scheduling, and structured generation. The PagedAttention paper describes virtual-memory-inspired KV-cache management to reduce fragmentation and duplication.
These papers cover techniques for managing KV-cache memory and improving LLM serving throughput and utilization.
Building Diffusion Language Models
A Kuleshov Group blog post introduces diffusion language models and research behind recent open-source models. It covers masking diffusion, iterative refinement, post-training, and variable-length generation.
Engineers can use it to understand the techniques that make diffusion language models work.
- Attention and model architecturePost on X
Qwen 3.5 and Nemotron 3 use different hybrid architectures
The post describes Qwen 3.5 as a Gated Delta Net hybrid and linear-attention variant, and Nemotron 3 as a Mamba hybrid. It connects linear attention and state-space models as forms of convergent evolution.
The comparison highlights architectural approaches beyond standard QKV attention.
Mechanistic Data Attribution Traces Training Data Behind LLM Circuits
The Mechanistic Data Attribution framework uses influence functions to link interpretable LLM units to training samples. The authors report that changing a few high-influence samples affects which heads emerge, and that interventions on induction heads alter in-context learning.
It connects training data to the emergence and behavior of interpretable attention heads.
HiLS learns chunk selection for sparse attention
The paper proposes Hierarchical Landmark Sparse (HiLS) Attention, a chunk-wise sparse attention method that learns chunk selection end-to-end under the language-modeling loss. The post says chunks are scored using aggregated scores from landmark tokens.
Engineers exploring long-context models can assess an approach to selecting chunks for sparse attention.
HOLA combines recurrent state with a bounded KV cache
The post describes HOLA, which pairs a recurrent state with a small KV cache for exact memory, routing tokens with high prediction error to the cache. It reports a 16.1% perplexity reduction and robust retrieval at 32k tokens.
The design explores whether a bounded exact-memory buffer can improve long-context retrieval in linear-attention models.
- Attention and model architecturePost on X
Functional Attention uses learned bases to reduce attention cost
Functional Attention represents query and key/value fields with learned bases, then relates them through a compact matrix instead of an all-pairs score table. The paper reports linear cost in the number of points and results across PDE, point-cloud segmentation, and regression tasks.
Engineers working on long sequences or operator learning can assess an alternative to quadratic all-pairs attention.
- Attention and model architecturePost on X
HOLA Adds a 64-Token Cache to Linear Attention
The post says HOLA uses a parameter-free token-surprise metric and a 64-token cache to address catastrophic forgetting in linear attention, claiming it beats full-attention baselines.
The approach could matter to engineers balancing inference cost, memory recall, and cache size.
FlowTracer Traces Attention-Based Information Flow for RL
The paper introduces FlowTracer, which traces information propagation through LLM reasoning to identify influential tokens and uses those signals to target reinforcement-learning credit assignment.
It offers a way to allocate RL training signals according to tokens that contribute to answer formation.
HOLA Adds a Bounded Exact KV Cache to Linear Attention
HOLA combines a delta-rule recurrent state with a bounded exact key-value cache. The paper reports improved Wikitext perplexity and robust RULER needle recall up to 32k tokens.
The design offers a way to improve long-range recall while retaining linear attention’s compressed state.
- Attention and model architecturePost on X
A Reading List on Transformers, Scaling, and Fine-Tuning
The post lists ten papers and works, in a suggested order, covering transformers, scaling, few-shot learning, RLHF, LoRA, FlashAttention, chain-of-thought, and DPO.
It offers engineers a structured path through influential LLM architecture and training concepts.
A Study of Shared QKV Projections in Transformers
The paper systematically evaluates Transformer variants that share key and value, query and key, or all three projections. The post reports that setting K equal to V was too much for the author's small ternary LLMs.
Projection sharing can reduce the number of distinct attention projections, and the results may inform architecture choices.
- Attention and model architecturePost on X
Video explains the transformer attention pipeline
The video covers Q/K/V, attention scores, and encoder blocks. The post describes attention as matching each token’s query against sequence keys to weight value vectors.
A concise walkthrough of the core operations behind transformer attention may help engineers build intuition for model architecture.
- Attention and model architecturePost on X
Replacing Transformer Attention Heads with Synthesized Programs
The post describes synthesizing Python programs from attention maps to reproduce individual heads’ token-to-token patterns. On GPT-2, TinyLlama, and Llama-3B, the best programs reached 69–79% mean IoU; replacing up to 30–40% of heads reportedly caused little QA decline.
This explores whether some attention heads can be approximated by explicit, swappable programs rather than neural computations.
- Attention and model architecturePost on X
HydraHead Mixes Full and Linear Attention by Head
HydraHead is an attention hybridization architecture that combines Full Attention and Linear Attention at the head level. The post says it is motivated by mechanistic interpretability and aims to support more efficient long-context models.
The head-level design offers an alternative granularity for combining attention mechanisms in long-context models.
- Attention and model architecturePost on X
Explaining Transformer Attention Heads with Python Programs
A paper introduces an automated interpretability technique that explains attention heads with Python programs. Replacing about 40% of attention patterns in Llama-3B with program outputs barely affects task performance.
The technique could help engineers identify attention patterns that can be simplified or inform architectural changes.
Explaining Transformer Attention Heads with Programs
The paper proposes approximating deep-network components with executable programs, focusing on attention heads in transformer language models. It computes attention matrices on randomly selected training examples and prompts a pretrained language model.
Programmatic approximations of attention heads could help engineers study how transformer components behave.
How reasoning unlocks knowledge stored in LLMs
A Google Research blog studies how reasoning helps LLMs recover knowledge stored in their weights.
It examines how reasoning can affect access to an LLM’s parametric knowledge.
Geometric Structures Behind Arithmetic in LLMs
The paper analyzes residual-stream representations during multi-operand addition and identifies the Iso-Raw-Sum Trajectory, where semantic digits anchor representations and continuous carry fibers modulate them. It proposes the Noisy Quantization Model to explain arithmetic errors.
It offers a geometric account of how LLM representations encode carries and where arithmetic errors may arise.
Build with AgentLog
Send your own newsletter
CommsHarbor keeps contacts, consent and one-click unsubscribe together.
Free workspace
Open CommsHarbor