Skip to content
EN

Attention and model architecture

Transformers, attention variants, state-space models and the ideas behind new architectures.

156 links, newest first.

Get the weekly briefing

The best new links of the topics you pick, summarized with the source. At most one email a week.

Topics: Attention and model architecture

Before the first issue we email you to confirm; leaving takes one click. Sent with CommsHarbor. Privacy

  1. iLLaDA Trains an 8B Diffusion Language Model

    The paper describes training an 8B Transformer with bidirectional masked diffusion on 12T tokens, followed by diffusion-based instruction tuning. It reports improved benchmark scores over LLaDA, while iLLaDA-Instruct still trails Qwen2.5 Instruct without RL alignment.

    It offers details on diffusion-model training and instruction tuning that engineers can compare with autoregressive approaches.

  2. iLLaDA: an 8B masked diffusion language model

    iLLaDA is an 8B masked diffusion language model trained from scratch with fully bidirectional attention. It retains the masked diffusion objective during pre-training and supervised fine-tuning.

    It offers engineers a concrete example of a language model architecture using bidirectional attention and a diffusion objective throughout training.

  3. FuncAttn Applies Functional Attention to Continuous Data

    The post describes FuncAttn, which represents attention between function spaces with structured linear operators instead of token-level softmax matching. It claims the method is compact and resolution-invariant, with results on PDE solving, 3D segmentation, and regression.

    The approach may be relevant to engineers working with continuous data and scientific machine learning.

  4. The Principles of Deep Learning Theory

    This book develops an effective-theory approach to deep neural networks, using layer-to-layer equations and nonlinear learning dynamics to describe trained-network outputs.

    It offers a theoretical framework for analyzing how neural network structure and training shape model behavior.

  5. The Transformer Cookbook explores algorithms in transformer weights

    The Transformer Cookbook introduces “hardcoding” algorithms such as addition, lookup, and branching into transformer weights, following the RASP paper.

    It offers engineers a way to think about implementing algorithmic behavior within transformers.

  6. OpenMythos proposes a recurrent-depth Transformer architecture

    OpenMythos is an open-source theoretical reconstruction of Claude Mythos. Its README describes a recurrent-depth Transformer with MoE routing, loop-index positional embeddings, and per-token ACT halting.

    The repository offers an inspectable implementation of architectural hypotheses involving recurrent depth, sparse experts, and adaptive halting.

  7. Chapter 8 introduces Transformer architecture

    Chapter 8 of *Speech and Language Processing* presents Transformer components including self-attention, Q/K/V vectors, multi-head attention, residual streams, and masking, building from an intuitive view to a matrix formulation.

    A step-by-step treatment can help engineers connect Transformer components to their underlying mathematical operations.

  8. NextLat Predicts Transformers’ Next Latent State

    Next-Latent Prediction (NextLat) is described as a self-supervised method that trains transformers to predict their next latent state. The post claims it enables compact world models and up to 3.3× faster inference through self-speculative decoding.

    The approach may interest engineers exploring alternative training objectives and faster transformer inference.

  9. Analyzing Efficient Attention in Hybrid Architectures

    The paper studies how sliding-window attention and recurrent sequence mixers shape capabilities in hybrid language models, examining scaling behavior, mechanisms, and architecture design.

    Its analysis may help engineers reason about the trade-offs of efficient attention modules in hybrid models.

  10. Dynamic Linear Attention adapts memory to token-level information

    Dynamic Linear Attention uses a State Information Score to add memory states at information transitions and merge adjacent low-information states, keeping a fixed cache. The post reports gains over Log Linear Attention across several benchmarks, with improved throughput and memory.

    Adaptive memory allocation could help engineers manage long-context quality and resource use.

  11. Fast KV Cache Compaction via Attention Matching

    The method builds compact key-value caches in latent space to preserve per-head attention outputs without slow end-to-end training. The post reports up to 50× compaction in seconds on some datasets, with minimal quality loss.

    Compact KV caches can reduce language-model memory use while preserving attention outputs.

  12. Lattice Deduction Transformers for Sudoku

    The post introduces an 800k-parameter looped transformer that it says reasons like a SAT solver and reached 100% on Sudoku-Extreme after 15 minutes of training.

    It describes a compact looped-transformer approach to reasoning on a structured puzzle task.

  13. Lecture Notes on Efficient LLM Deployment

    Lecture notes cover SmoothQuant, AWQ, INT4 inference kernels, pruning and sparsity, MoE, PagedAttention, FlashAttention, speculative decoding, and batching.

    The notes survey techniques and systems components used to improve LLM inference and serving.

  14. Stanford CS336 covers modern language model systems

    The 2026 edition of CS336 covers topics including MoEs, GPU tiling, kernels, RLHF, and data. The post says lectures appear on YouTube with about a two-week delay.

    Useful for engineers seeking a course spanning language model architecture and implementation topics.

  15. The Math Behind Query, Key, and Value in Attention

    The article explains attention using the formula softmax(Q × Kᵀ / √dₖ) × V and a step-by-step numeric example.

    It offers a concrete walkthrough of the computation at the core of transformer attention.

  16. Mamba-3 uses a fixed-size state for sequence modeling

    Mamba-3 is a non-Transformer language model architecture that compresses past context into a fixed-size state. The post reports decoding time per step is constant with sequence length and faster inference than a Transformer above 16k context on a 1.5B-parameter model.

    Engineers evaluating long-context inference can compare its state-space approach with Transformer KV caching.

  17. Two papers on selective attention and subquadratic limits

    The post points to two papers, one on selective attention and one on limitations of subquadratic methods. The author says they provide context for hurdles facing SubQ.

    Useful background for evaluating claims about sparse attention and subquadratic architectures.

  18. Triton and PyTorch Ops for Block Attention Residuals

    A GitHub repository providing Triton kernels and PyTorch ops for Block Attention Residuals (AttnRes).

    Engineers exploring AttnRes can review its kernel and operator implementation.

  19. microGPT Implemented in Dependency-Free C

    A GitHub project implements Karpathy’s microGPT in pure C. The post says it uses AVX2 intrinsics, fixed-width dot-product kernels, and a fast exponential approximation.

    A compact C implementation can help engineers inspect GPT training and inference while exploring low-level optimizations.

  20. DeepSeek-V4’s Shift to High-Rank MQA

    The post describes DeepSeek-V4 as using shared-KV MQA with 128 (64) query heads and 512-dimensional heads, contrasting it with MLA’s low-rank design and KV-cache focus.

    The comparison highlights how attention design can trade KV-cache efficiency for greater per-token computation and representation capacity.

Build with AgentLog

List your MCP, skill or plugin

Reach the engineers who read these briefings.

Sponsor AgentLog

Footer, sidebar or featured slot for 30 days.

From US$ 60

See the slots

Send your own newsletter

CommsHarbor keeps contacts, consent and one-click unsubscribe together.

Free workspace

Open CommsHarbor