Attention and model architecture
Transformers, attention variants, state-space models and the ideas behind new architectures.
156 links, newest first.
Google Research introduces Sequential Attention
Google Research presents Sequential Attention, which greedily selects important features or layers during training. The linked article describes it as aiming to make AI models leaner and faster without sacrificing accuracy.
Engineers exploring model efficiency can review an attention approach that selects features or layers during training.
- Attention and model architecturePost on X
Multi-Head LatentMoE with Head Parallelism
The post describes splitting each token into heads, distributing them across GPUs, then performing routing and expert work locally. It claims up to 1.61× speedup over standard MoE with expert parallelism and constant communication as expert count grows.
The design may reduce communication overhead and improve GPU balance in MoE systems.
- Attention and model architecturePost on X
AlphaEvolve research explores generalizable activation functions
The post says DeepMind is using AlphaEvolve to discover activation functions and describes this as part of broader research into improving the rate of novel architecture discovery.
Activation-function research may inform new model architectures and self-improving AI systems.
- Attention and model architectureRepository
Architecture and training choices discussed for nanochat
The post lists nanochat design choices including RoPE, parameter-free RMSNorm, squared ReLU, sliding-window attention, gated value embeddings, and learnable residual scalars. It also mentions split Muon/AdamW optimization.
It gives engineers a concise overview of architectural and optimization choices to investigate in the nanochat discussion.
- Attention and model architectureRepository
Scaled dot-product attention explained
The post describes how scaled dot-product attention uses queries, keys, and values to compare token relevance and weight information passed into each representation. It links to a notebook in a GitHub repository about implementing LLM and NLP systems from first principles.
A concise overview of the core attention mechanism, with a linked implementation notebook for further study.
- Attention and model architecturePost on X
CALM predicts vector chunks instead of next tokens
The post describes Continuous Autoregressive Language Models (CALM) as predicting “next-vectors” for chunks of meaning rather than individual word pieces. It claims this carries 4× more information per step and reduces training compute by 44%.
The proposed prediction unit is a potentially relevant alternative to token-by-token autoregression.
- Attention and model architecturePost on X
Universal Reasoning Model Uses Recurrence and Nonlinearity
The post describes research attributing Universal Transformer reasoning gains mainly to recurrent inductive bias and strong nonlinearity. The proposed URM adds ConvSwiGLU for local token mixing and truncated backpropagation through recurrent loops.
Its design choices offer engineers concrete ways to improve recurrent reasoning models and stabilize their training.
- Attention and model architecturePost on X
Kascade Reuses Key Patterns for Sparse Attention
The post describes Kascade, a training-free sparse-attention method that reuses key patterns across layers. It claims up to 4.1× faster generation and 2.2× faster prefill on H100 GPUs with little accuracy loss on long-context benchmarks.
Engineers working on long-context inference may find its reported speedups relevant.
Why Masked Diffusion Models Struggle with Token Ordering
An ICML 2025 Outstanding Paper studies why masked models underperform on generation benchmarks. It argues that training across arbitrary revealed-token subsets creates many computationally intractable subproblems, and reports uneven performance across masking patterns.
The findings can inform training and inference strategies for masked language models.
SEAL decomposes and checks knowledge-graph questions
SEAL is a framework for conversational question answering over knowledge graphs. The post describes extracting a minimal query core and using an agent to validate it against the database.
Its query decomposition and validation approach may help engineers address errors in multi-hop knowledge-graph question answering.
BEAVER computes sound probability bounds for LLM safety
BEAVER is a framework for computing deterministic, sound probability bounds on whether an LLM satisfies a given safety property. The paper contrasts these guarantees with sampling-based estimates.
Engineers evaluating LLM safety can use probability bounds rather than relying only on sampling-based estimates.
Post-training makes transformer attention sparse
The paper introduces constrained-loss sparsity regularization for transformer attention. On models up to 7B parameters, it reports retaining the original pretraining loss with about 0.4% of attention edges.
Engineers studying attention circuits can explore sparsity as a structural prior for interpretability.
Tensor Product Attention Uses Low-Rank Components to Shrink KV Caches
The paper proposes Tensor Product Attention (TPA), which uses tensor decompositions to represent queries, keys, and values with contextual low-rank components. It aims to reduce KV-cache memory overhead during inference.
Smaller KV caches could help engineers manage inference memory for language models with longer input sequences.
KArAt replaces softmax with learnable attention
The paper introduces Kolmogorov-Arnold Attention (KArAt), a learnable alternative to fixed softmax attention. The post reports mixed results across vision transformers and describes low-rank variants for reducing memory use.
It explores attention that may offer more inspectable interactions, while highlighting trade-offs across model sizes.
- Attention and model architecturePost on X
MoGA: Mixture-of-Groups Attention for Long Video Generation
The post names MoGA, a Mixture-of-Groups Attention approach for end-to-end long video generation.
Engineers working on attention or video-generation architectures may find the approach relevant.
- Attention and model architecturePost on X
Recursive Language Models Process Long Prompts Through a REPL
The post describes Recursive Language Models (RLMs), an inference strategy that lets LLMs decompose and recursively interact with very long prompts through a REPL. It reports benchmark results on OOLONG and BrowseComp-Plus.
The approach may offer an alternative to explicit retrieval for handling long-context tasks.
- Attention and model architecturePost on X
Circuit-Based Verification Predicts Reasoning Errors
The post describes Circuit-based Reasoning Verification (CRV), which uses attribution-graph structure to predict incorrect reasoning. On arithmetic tasks, it reports AUROC rising from 76.45 to 92.47 and FPR@95 falling from 63.33% to 37.09%.
It explores using model-internal circuit structure to detect reasoning failures.
- Attention and model architecturePost on X
SwiReasoning switches between latent and explicit reasoning
The post describes SwiReasoning, which uses predictive entropy to switch between latent reasoning with soft embeddings and explicit Chain-of-Thought. It claims a peak 6.78× token-efficiency gain and an average 56–79% gain on Qwen3-8B models.
The confidence-triggered switching approach may interest engineers exploring ways to reduce reasoning-token costs.
- Attention and model architecturePost on X
A study compares attention levels in 13 LLMs
The author reports training 13 LLMs with varying proportions of attention and DeltaNet linear attention. In their experiments, 17% attention—two of 12 layers—performed best.
The result may inform how engineers balance full attention and linear attention in model architectures.
- Attention and model architecturePost on X
A Three-Stage Theory of Information Flow in LLMs
The post describes a theory linking massive activations to attention sinks and compression valleys in LLMs, and proposes three stages of information flow.
It may help engineers reason about how activations relate to attention and information flow in LLMs.
Build with AgentLog
Send your own newsletter
CommsHarbor keeps contacts, consent and one-click unsubscribe together.
Free workspace
Open CommsHarbor

