Skip to content
EN

Attention and model architecture

Transformers, attention variants, state-space models and the ideas behind new architectures.

156 links, newest first.

Get the weekly briefing

The best new links of the topics you pick, summarized with the source. At most one email a week.

Topics: Attention and model architecture

Before the first issue we email you to confirm; leaving takes one click. Sent with CommsHarbor. Privacy

  1. A Faithful Implementation of DeepSeek-V4 Compressed Sparse Attention

    The linked page presents a faithful implementation of Compressed Sparse Attention (CSA) from the DeepSeek-V4 paper.

    It may help engineers examining an implementation of an attention variant from a recent model paper.

  2. Study examines knowledge capacity in attention and MLPs

    The work analyzes trade-offs in how attention and MLP matrices store knowledge, using a needle-in-a-haystack model. It highlights limits from finite samples and compute, as well as noise in long contexts.

    It frames how data, compute, and context noise can constrain model knowledge storage.

  3. Looped Transformers as Energy-Based Model Inference

    The post relates residual updates in a looped transformer to gradient descent on an energy function, when the block equals the negative energy gradient. It notes that a generic transformer block does not automatically satisfy this condition.

    This framing connects weight reuse in looped transformers with iterative optimization, while highlighting a condition the block must meet.

  4. Latent Reasoning Tuning uses compact internal representations

    Latent Reasoning Tuning (LRT) trains models to reason with compact internal representations rather than long text chains. The post claims it outperforms other efficient reasoning methods and a Qwen3 hybrid framework on math and general benchmarks.

    The approach explores reducing the cost of reasoning by avoiding long generated chains of thought.

  5. HybridGen Enables CPU–GPU Attention for Long-Context Inference

    The paper proposes HybridGen, a hybrid attention framework for efficient long-context LLM inference. It enables CPU–GPU collaborative attention on systems with CXL memory.

    Engineers working on long-context inference can assess an approach to coordinating CPU and GPU attention with CXL memory.

  6. CliffordNet Uses Clifford Algebra for Neural Network Interactions

    The post describes CliffordNet as an architecture based on the geometric product, combining inner-product and geometric interactions without attention or mixer layers. It reports 77.82% accuracy on CIFAR-100 with 1.4M parameters.

    It presents an alternative to attention-based and mixer architectures that engineers can evaluate against the linked paper.

  7. LACE Enables Cross-Thread Attention Between Reasoning Paths

    LACE is a framework that lets parallel reasoning paths share intermediate insights and correct one another through cross-thread attention. It repurposes the model architecture for coordinated reasoning.

    Engineers can explore an architecture for reducing redundant parallel reasoning and sharing useful intermediate results.

  8. Looped Transformers as Programmable Computers

    The paper presents a framework that programs transformer networks with specific weights and places them in a loop. It shows how a constant number of encoder layers can emulate basic computing blocks.

    It offers a construction for understanding how looped transformers can implement programmable computation.

  9. Mixture-of-Depths Attention connects attention across layers

    The paper introduces MoDA, which lets each attention head attend to sequence key-value pairs at the current layer and depth key-value pairs from preceding layers. It targets signal degradation in deeper LLMs.

    Engineers exploring attention architectures can assess a way to reuse features from earlier layers.

  10. A guide to attention mechanisms in Transformers and LLMs

    The article explains more than 13 attention mechanisms, including self-attention, FlashAttention, GQA, MLA, and sliding-window attention. It describes them as approaches that shape how Transformers and LLMs process context and memory.

    A concise overview can help engineers compare attention variants when evaluating model architectures.

  11. Training Compute and Memory Costs of Loop Models

    A blog analyzes loop-model training through a seven-variant ablation chain, focusing on how gradient structure affects FLOPs and memory. It includes formulas and toy profiler checks.

    Useful for estimating the training costs of loop architectures and understanding how gradient structure changes them.

  12. Interleaved Head Attention Explores Communication Across Heads

    The post introduces Interleaved Head Attention (IHA), motivated by the idea that attention heads may benefit from communicating rather than remaining isolated. It frames the question from a one-layer reasoning perspective.

    It raises an architectural question about how attention heads can communicate within a layer.

  13. DeepCrossAttention uses input-dependent layer mixing

    DeepCrossAttention (DCA) introduces learnable, input-dependent weights to dynamically combine Transformer layer outputs, as an alternative to summing them through traditional residual connections.

    It offers an architecture approach to residual learning that engineers can examine for richer interactions across Transformer layers.

  14. Why Low-Precision Flash Attention Training Can Fail

    The paper analyzes catastrophic loss explosions during low-precision training with Flash Attention. It attributes the failure to similar internal attention data patterns and biased rounding errors that accumulate.

    Understanding these failure mechanisms can help engineers diagnose instability in low-precision Transformer training.

  15. Exclusive Self-Attention Removes Each Token’s Own Value

    The paper introduces exclusive self-attention, which constrains attention to capture information orthogonal to each token’s own value vector. On standard language modeling, it reports better sequence-modeling performance than self-attention across model sizes up to 2.7B parameters, with larger…

    Engineers can assess a simple attention modification that the paper reports improves language-modeling performance.

  16. FlashAttention-4 targets NVIDIA Blackwell GPUs

    FlashAttention-4 uses new pipelines, reworked computations, and memory optimizations for Blackwell GPUs. The post reports speedups over cuDNN 9.13 and Triton on B200 GPUs.

    Its reported performance and compile-time changes may matter when optimizing attention workloads on Blackwell GPUs.

  17. Kascade reuses sparse attention across transformer layers

    Kascade computes full attention in selected reference layers and reuses the results in intervening layers. The method is designed to reduce long-context inference work without retraining or changing model weights.

    It offers an inference-time approach to reducing attention computation in long-context models.

  18. Three Approaches to On-Policy Self-Distillation

    The article examines OPSD, SDFT, and SDPO, three on-policy self-distillation methods that use denser feedback. It covers their benchmarks and limitations.

    The comparison helps engineers understand how these methods use feedback to improve model reasoning.

  19. Graph Attention Networks Learn Which Neighbors Matter

    The ICLR 2018 paper introduces Graph Attention Networks, which use attention to weight information from neighboring nodes. A shared neural network computes attention weights for connected node pairs.

    It offers an attention-based approach to aggregating graph data that engineers can compare with other graph neural network architectures.

  20. ROME Uses Chunk-Level Credit for Agent Training

    The post describes ROME as a 3B-parameter model trained with IPA, which assigns reward at the chunk level. It lists structured-code pretraining, error-masked SFT, and chunk-level RL optimization.

    Chunk-level reward assignment offers an alternative to token-level or trajectory-level credit for tool-using agents.

Build with AgentLog

List your MCP, skill or plugin

Reach the engineers who read these briefings.

Sponsor AgentLog

Footer, sidebar or featured slot for 30 days.

From US$ 60

See the slots

Send your own newsletter

CommsHarbor keeps contacts, consent and one-click unsubscribe together.

Free workspace

Open CommsHarbor