Attention and model architecture
Transformers, attention variants, state-space models and the ideas behind new architectures.
156 links, newest first.
A Faithful Implementation of DeepSeek-V4 Compressed Sparse Attention
The linked page presents a faithful implementation of Compressed Sparse Attention (CSA) from the DeepSeek-V4 paper.
It may help engineers examining an implementation of an attention variant from a recent model paper.
- Attention and model architecturePost on X
Study examines knowledge capacity in attention and MLPs
The work analyzes trade-offs in how attention and MLP matrices store knowledge, using a needle-in-a-haystack model. It highlights limits from finite samples and compute, as well as noise in long contexts.
It frames how data, compute, and context noise can constrain model knowledge storage.
- Attention and model architecturePost on X
Looped Transformers as Energy-Based Model Inference
The post relates residual updates in a looped transformer to gradient descent on an energy function, when the block equals the negative energy gradient. It notes that a generic transformer block does not automatically satisfy this condition.
This framing connects weight reuse in looped transformers with iterative optimization, while highlighting a condition the block must meet.
- Attention and model architectureRepository
Latent Reasoning Tuning uses compact internal representations
Latent Reasoning Tuning (LRT) trains models to reason with compact internal representations rather than long text chains. The post claims it outperforms other efficient reasoning methods and a Qwen3 hybrid framework on math and general benchmarks.
The approach explores reducing the cost of reasoning by avoiding long generated chains of thought.
HybridGen Enables CPU–GPU Attention for Long-Context Inference
The paper proposes HybridGen, a hybrid attention framework for efficient long-context LLM inference. It enables CPU–GPU collaborative attention on systems with CXL memory.
Engineers working on long-context inference can assess an approach to coordinating CPU and GPU attention with CXL memory.
CliffordNet Uses Clifford Algebra for Neural Network Interactions
The post describes CliffordNet as an architecture based on the geometric product, combining inner-product and geometric interactions without attention or mixer layers. It reports 77.82% accuracy on CIFAR-100 with 1.4M parameters.
It presents an alternative to attention-based and mixer architectures that engineers can evaluate against the linked paper.
LACE Enables Cross-Thread Attention Between Reasoning Paths
LACE is a framework that lets parallel reasoning paths share intermediate insights and correct one another through cross-thread attention. It repurposes the model architecture for coordinated reasoning.
Engineers can explore an architecture for reducing redundant parallel reasoning and sharing useful intermediate results.
Looped Transformers as Programmable Computers
The paper presents a framework that programs transformer networks with specific weights and places them in a loop. It shows how a constant number of encoder layers can emulate basic computing blocks.
It offers a construction for understanding how looped transformers can implement programmable computation.
Mixture-of-Depths Attention connects attention across layers
The paper introduces MoDA, which lets each attention head attend to sequence key-value pairs at the current layer and depth key-value pairs from preceding layers. It targets signal degradation in deeper LLMs.
Engineers exploring attention architectures can assess a way to reuse features from earlier layers.
A guide to attention mechanisms in Transformers and LLMs
The article explains more than 13 attention mechanisms, including self-attention, FlashAttention, GQA, MLA, and sliding-window attention. It describes them as approaches that shape how Transformers and LLMs process context and memory.
A concise overview can help engineers compare attention variants when evaluating model architectures.
Training Compute and Memory Costs of Loop Models
A blog analyzes loop-model training through a seven-variant ablation chain, focusing on how gradient structure affects FLOPs and memory. It includes formulas and toy profiler checks.
Useful for estimating the training costs of loop architectures and understanding how gradient structure changes them.
Interleaved Head Attention Explores Communication Across Heads
The post introduces Interleaved Head Attention (IHA), motivated by the idea that attention heads may benefit from communicating rather than remaining isolated. It frames the question from a one-layer reasoning perspective.
It raises an architectural question about how attention heads can communicate within a layer.
DeepCrossAttention uses input-dependent layer mixing
DeepCrossAttention (DCA) introduces learnable, input-dependent weights to dynamically combine Transformer layer outputs, as an alternative to summing them through traditional residual connections.
It offers an architecture approach to residual learning that engineers can examine for richer interactions across Transformer layers.
Why Low-Precision Flash Attention Training Can Fail
The paper analyzes catastrophic loss explosions during low-precision training with Flash Attention. It attributes the failure to similar internal attention data patterns and biased rounding errors that accumulate.
Understanding these failure mechanisms can help engineers diagnose instability in low-precision Transformer training.
Exclusive Self-Attention Removes Each Token’s Own Value
The paper introduces exclusive self-attention, which constrains attention to capture information orthogonal to each token’s own value vector. On standard language modeling, it reports better sequence-modeling performance than self-attention across model sizes up to 2.7B parameters, with larger…
Engineers can assess a simple attention modification that the paper reports improves language-modeling performance.
- Attention and model architecturePost on X
FlashAttention-4 targets NVIDIA Blackwell GPUs
FlashAttention-4 uses new pipelines, reworked computations, and memory optimizations for Blackwell GPUs. The post reports speedups over cuDNN 9.13 and Triton on B200 GPUs.
Its reported performance and compile-time changes may matter when optimizing attention workloads on Blackwell GPUs.
Kascade reuses sparse attention across transformer layers
Kascade computes full attention in selected reference layers and reuses the results in intervening layers. The method is designed to reduce long-context inference work without retraining or changing model weights.
It offers an inference-time approach to reducing attention computation in long-context models.
Three Approaches to On-Policy Self-Distillation
The article examines OPSD, SDFT, and SDPO, three on-policy self-distillation methods that use denser feedback. It covers their benchmarks and limitations.
The comparison helps engineers understand how these methods use feedback to improve model reasoning.
Graph Attention Networks Learn Which Neighbors Matter
The ICLR 2018 paper introduces Graph Attention Networks, which use attention to weight information from neighboring nodes. A shared neural network computes attention weights for connected node pairs.
It offers an attention-based approach to aggregating graph data that engineers can compare with other graph neural network architectures.
- Attention and model architecturePost on X
ROME Uses Chunk-Level Credit for Agent Training
The post describes ROME as a 3B-parameter model trained with IPA, which assigns reward at the chunk level. It lists structured-code pretraining, error-masked SFT, and chunk-level RL optimization.
Chunk-level reward assignment offers an alternative to token-level or trajectory-level credit for tool-using agents.
Build with AgentLog
Send your own newsletter
CommsHarbor keeps contacts, consent and one-click unsubscribe together.
Free workspace
Open CommsHarbor


