Skip to content
EN

Attention and model architecture

Transformers, attention variants, state-space models and the ideas behind new architectures.

156 links, newest first.

Get the weekly briefing

The best new links of the topics you pick, summarized with the source. At most one email a week.

Topics: Attention and model architecture

Before the first issue we email you to confirm; leaving takes one click. Sent with CommsHarbor. Privacy

  1. Google Research introduces Sequential Attention

    Google Research presents Sequential Attention, which greedily selects important features or layers during training. The linked article describes it as aiming to make AI models leaner and faster without sacrificing accuracy.

    Engineers exploring model efficiency can review an attention approach that selects features or layers during training.

  2. Multi-Head LatentMoE with Head Parallelism

    The post describes splitting each token into heads, distributing them across GPUs, then performing routing and expert work locally. It claims up to 1.61× speedup over standard MoE with expert parallelism and constant communication as expert count grows.

    The design may reduce communication overhead and improve GPU balance in MoE systems.

  3. AlphaEvolve research explores generalizable activation functions

    The post says DeepMind is using AlphaEvolve to discover activation functions and describes this as part of broader research into improving the rate of novel architecture discovery.

    Activation-function research may inform new model architectures and self-improving AI systems.

  4. Architecture and training choices discussed for nanochat

    The post lists nanochat design choices including RoPE, parameter-free RMSNorm, squared ReLU, sliding-window attention, gated value embeddings, and learnable residual scalars. It also mentions split Muon/AdamW optimization.

    It gives engineers a concise overview of architectural and optimization choices to investigate in the nanochat discussion.

  5. Scaled dot-product attention explained

    The post describes how scaled dot-product attention uses queries, keys, and values to compare token relevance and weight information passed into each representation. It links to a notebook in a GitHub repository about implementing LLM and NLP systems from first principles.

    A concise overview of the core attention mechanism, with a linked implementation notebook for further study.

  6. CALM predicts vector chunks instead of next tokens

    The post describes Continuous Autoregressive Language Models (CALM) as predicting “next-vectors” for chunks of meaning rather than individual word pieces. It claims this carries 4× more information per step and reduces training compute by 44%.

    The proposed prediction unit is a potentially relevant alternative to token-by-token autoregression.

  7. Universal Reasoning Model Uses Recurrence and Nonlinearity

    The post describes research attributing Universal Transformer reasoning gains mainly to recurrent inductive bias and strong nonlinearity. The proposed URM adds ConvSwiGLU for local token mixing and truncated backpropagation through recurrent loops.

    Its design choices offer engineers concrete ways to improve recurrent reasoning models and stabilize their training.

  8. Kascade Reuses Key Patterns for Sparse Attention

    The post describes Kascade, a training-free sparse-attention method that reuses key patterns across layers. It claims up to 4.1× faster generation and 2.2× faster prefill on H100 GPUs with little accuracy loss on long-context benchmarks.

    Engineers working on long-context inference may find its reported speedups relevant.

  9. Why Masked Diffusion Models Struggle with Token Ordering

    An ICML 2025 Outstanding Paper studies why masked models underperform on generation benchmarks. It argues that training across arbitrary revealed-token subsets creates many computationally intractable subproblems, and reports uneven performance across masking patterns.

    The findings can inform training and inference strategies for masked language models.

  10. SEAL decomposes and checks knowledge-graph questions

    SEAL is a framework for conversational question answering over knowledge graphs. The post describes extracting a minimal query core and using an agent to validate it against the database.

    Its query decomposition and validation approach may help engineers address errors in multi-hop knowledge-graph question answering.

  11. BEAVER computes sound probability bounds for LLM safety

    BEAVER is a framework for computing deterministic, sound probability bounds on whether an LLM satisfies a given safety property. The paper contrasts these guarantees with sampling-based estimates.

    Engineers evaluating LLM safety can use probability bounds rather than relying only on sampling-based estimates.

  12. Post-training makes transformer attention sparse

    The paper introduces constrained-loss sparsity regularization for transformer attention. On models up to 7B parameters, it reports retaining the original pretraining loss with about 0.4% of attention edges.

    Engineers studying attention circuits can explore sparsity as a structural prior for interpretability.

  13. Tensor Product Attention Uses Low-Rank Components to Shrink KV Caches

    The paper proposes Tensor Product Attention (TPA), which uses tensor decompositions to represent queries, keys, and values with contextual low-rank components. It aims to reduce KV-cache memory overhead during inference.

    Smaller KV caches could help engineers manage inference memory for language models with longer input sequences.

  14. KArAt replaces softmax with learnable attention

    The paper introduces Kolmogorov-Arnold Attention (KArAt), a learnable alternative to fixed softmax attention. The post reports mixed results across vision transformers and describes low-rank variants for reducing memory use.

    It explores attention that may offer more inspectable interactions, while highlighting trade-offs across model sizes.

  15. MoGA: Mixture-of-Groups Attention for Long Video Generation

    The post names MoGA, a Mixture-of-Groups Attention approach for end-to-end long video generation.

    Engineers working on attention or video-generation architectures may find the approach relevant.

  16. Recursive Language Models Process Long Prompts Through a REPL

    The post describes Recursive Language Models (RLMs), an inference strategy that lets LLMs decompose and recursively interact with very long prompts through a REPL. It reports benchmark results on OOLONG and BrowseComp-Plus.

    The approach may offer an alternative to explicit retrieval for handling long-context tasks.

  17. Circuit-Based Verification Predicts Reasoning Errors

    The post describes Circuit-based Reasoning Verification (CRV), which uses attribution-graph structure to predict incorrect reasoning. On arithmetic tasks, it reports AUROC rising from 76.45 to 92.47 and FPR@95 falling from 63.33% to 37.09%.

    It explores using model-internal circuit structure to detect reasoning failures.

  18. SwiReasoning switches between latent and explicit reasoning

    The post describes SwiReasoning, which uses predictive entropy to switch between latent reasoning with soft embeddings and explicit Chain-of-Thought. It claims a peak 6.78× token-efficiency gain and an average 56–79% gain on Qwen3-8B models.

    The confidence-triggered switching approach may interest engineers exploring ways to reduce reasoning-token costs.

  19. A study compares attention levels in 13 LLMs

    The author reports training 13 LLMs with varying proportions of attention and DeltaNet linear attention. In their experiments, 17% attention—two of 12 layers—performed best.

    The result may inform how engineers balance full attention and linear attention in model architectures.

  20. A Three-Stage Theory of Information Flow in LLMs

    The post describes a theory linking massive activations to attention sinks and compression valleys in LLMs, and proposes three stages of information flow.

    It may help engineers reason about how activations relate to attention and information flow in LLMs.

Build with AgentLog

List your MCP, skill or plugin

Reach the engineers who read these briefings.

Sponsor AgentLog

Footer, sidebar or featured slot for 30 days.

From US$ 60

See the slots

Send your own newsletter

CommsHarbor keeps contacts, consent and one-click unsubscribe together.

Free workspace

Open CommsHarbor