Skip to content
EN

Attention and model architecture

Transformers, attention variants, state-space models and the ideas behind new architectures.

156 links, newest first.

Get the weekly briefing

The best new links of the topics you pick, summarized with the source. At most one email a week.

Topics: Attention and model architecture

Before the first issue we email you to confirm; leaving takes one click. Sent with CommsHarbor. Privacy

  1. Xeno-Interpretability and Model-Native Representations

    The paper asks whether language models use internal distinctions that lack adequate human concepts. It calls these structures xeno-representations and studies them as xeno-interpretability.

    It examines a possible limit of interpreting model representations through familiar human concepts.

  2. mini-AGI explores continual learning on 8GB VRAM

    The mini-AGI repository describes a model trained from scratch on an 8GB VRAM laptop using a batch-1 stream of data. The post says it pages weights from disk and adjusts learning rates to reduce forgetting.

    It offers an implementation to examine for online learning and memory-constrained model training.

  3. On-Demand Attention Selects When to Use Full Context

    On-Demand Attention trains a lightweight head to invoke full context when needed. The post says it speeds up long-context decoding while recovering most performance lost under local attention.

    It offers an approach to balancing long-context decoding efficiency and performance.

  4. Gated Linear Attention combines linear attention with data-dependent gates

    The paper presents Gated Linear Attention, a linear-attention transformer formulation that can also be expressed as an RNN with matrix-valued hidden states. It describes a hardware-efficient training algorithm for linear attention.

    Engineers can assess an approach to linear-time inference and hardware-efficient training as alternatives to softmax attention.

  5. Jamba combines Transformer and Mamba layers

    Jamba is a base language model with a hybrid Transformer-Mamba mixture-of-experts architecture. It interleaves Transformer and Mamba blocks, with MoE in some layers to increase capacity while managing active parameter usage.

    The paper describes a hybrid architecture and configurable trade-offs for model capacity and resource use.

  6. RWKV combines parallel training with recurrent inference

    The RWKV paper proposes an architecture combining Transformer-style parallel training with RNN-style recurrent inference. Its abstract describes linear memory and computational scaling with sequence length.

    Engineers evaluating alternatives for long-context workloads can compare RWKV’s recurrent inference design with Transformer architectures.

  7. Flow Models as Recurrent Models for Reasoning

    The work turns a flow model into a recurrent model for tasks such as Sudoku. It feeds predictions back as inputs for iterative refinement, using a local loss at each pass and no backpropagation through time.

    The approach offers an alternative way to refine predictions across multiple passes without backpropagation through time.

  8. Block-Recurrent Transformers Process Sequences with Linear Complexity

    The paper introduces a transformer layer applied recurrently along a sequence, with linear complexity in sequence length. Its recurrent cell operates on token blocks and uses self-attention and cross-attention.

    Block-level recurrence offers an alternative architecture for processing long sequences with linear sequence-length complexity.

  9. Universal Transformers Reuse a Transformer Layer Recurrently

    The post describes Universal Transformers as iteratively applying the same Transformer layer instead of stacking distinct layers. The linked paper discusses sequence-modeling architectures and the trade-off between sequential and parallel computation.

    Layer reuse is an architectural alternative to stacking layers that engineers can evaluate for sequence models.

  10. Looped Transformers as Programmable Computers

    The paper presents a framework for programming Transformer networks with specific weights and placing them in a loop to act as universal computers. It describes encoder layers that emulate basic computing blocks, including conditional branches and program counters.

    It explores how recurrent use of Transformer layers can support algorithmic computation.

  11. Thinking with Looped Flows Uses Local Denoising Objectives

    Looped flows train recurrent models with local denoising objectives to improve multi-step computation. The post reports better results than prior looped models across six reasoning benchmarks, including 58.8% accuracy on ARC-AGI-1.

    Engineers exploring recurrent architectures can compare this training approach and its reported reasoning benchmark results.

  12. GOAT introduces trainable priors for attention

    The paper frames attention as Entropic Optimal Transport and introduces GOAT, which replaces the implicit uniform prior with a learnable, continuous prior. The method is compatible with optimized kernels such as FlashAttention.

    Engineers can assess an attention variant that changes the prior while retaining compatibility with optimized kernels.

  13. Softmax Linear Attention restores competition with linear complexity

    The paper proposes Softmax Linear Attention (SLA), applying softmax at the head level to restore global competition in linear attention. It aims to retain linear-time complexity while improving selection of relevant information.

    Engineers evaluating attention variants can compare SLA’s approach to efficient attention with standard softmax normalization.

  14. Tucker Attention generalizes MHA, GQA, and MLA

    The paper presents Tucker Attention, a tensor-factorization approach that frames MHA, GQA, and MLA as approximate attention methods based on specialized low-rank factorizations.

    It offers a unified low-rank perspective for engineers evaluating attention variants.

  15. KATA uses spherical packing for linear attention recall

    The paper frames associative recall in linear attention as a spherical-packing problem and introduces Kernelized Linear Attention Activations (KATA), with feature maps derived from a self-dual homogeneous cone.

    It explores a geometric approach to improving associative-recall capacity in linear attention.

  16. SAS trains a selector inside Softmax to filter context

    The post describes a SAS paper that trains an end-to-end selector inside Softmax to filter context for attention on compute-constrained hardware.

    Engineers working on attention efficiency can examine the proposed context-filtering approach.

  17. Critical Attention Scaling in Long-Context Transformers

    The paper analyzes rank-collapse in attention as context length grows and studies rescaling attention scores by a polylogarithmic factor. The post says it justifies logarithmic scaling for long contexts.

    It offers theoretical analysis of attention scaling choices for long-context models.

  18. PC-ALM trains deep residual MLPs with local dynamics

    PC-ALM is a local alternative to backpropagation that uses augmented Lagrangian predictive coding. It trains residual MLPs up to 1000 layers, with performance described as nearly matching backpropagation.

    It explores how local learning dynamics can propagate training signals through very deep networks without backpropagation.

  19. Simple Attention Sparsification Optimizes KV Block Selection

    Tencent’s Simple Attention Sparsification lets language-modeling loss directly optimize which KV blocks each query attends to.

    It describes a way to train sparse attention by learning query-to-KV-block selection.

  20. Token Retention as a Zero-Training KV Cache Alternative

    The post claims that retaining four initial tokens and a 64-token window, without training, matches or beats distilled linear-attention models on almost every benchmark.

    It highlights a simple token-retention strategy that may reduce reliance on the KV cache without post-training.

Build with AgentLog

List your MCP, skill or plugin

Reach the engineers who read these briefings.

Sponsor AgentLog

Footer, sidebar or featured slot for 30 days.

From US$ 60

See the slots

Send your own newsletter

CommsHarbor keeps contacts, consent and one-click unsubscribe together.

Free workspace

Open CommsHarbor