Skip to content
EN

Attention and model architecture

Transformers, attention variants, state-space models and the ideas behind new architectures.

156 links, newest first.

Get the weekly briefing

The best new links of the topics you pick, summarized with the source. At most one email a week.

Topics: Attention and model architecture

Before the first issue we email you to confirm; leaving takes one click. Sent with CommsHarbor. Privacy

  1. Attention Sinks and Compression Valleys in LLMs

    Experiments across models from 410M to 120B parameters report that extreme middle-layer activation norms at the beginning-of-sequence token coincide with compression valleys and attention sinks. The authors propose three phases of Transformer computation: broad mixing, limited mixing, and…

    The proposed phase model may help engineers reason about how information mixing changes across Transformer layers.

  2. Artificial Hippocampus Networks for Long-Context LLMs

    A post says ByteDance released Artificial Hippocampus Networks (AHN), an architecture for long-context LLMs that continuously compresses information outside the context window.

    Its approach to retaining out-of-window information may interest engineers exploring memory-efficient long-context architectures.

  3. CCA Runs Attention in a Compressed Latent Space

    The post describes Compressed Contextual Attention (CCA), which computes attention in a smaller latent space and applies RoPE there. It reports reduced KV-cache, parameter, and FLOP costs, plus faster prefill and backward passes in H100 tests.

    CCA offers an attention design engineers can evaluate when balancing compute, memory, and model quality.

  4. Compressed Convolutional Attention for Efficient Attention

    The paper introduces Compressed Convolutional Attention (CCA), which down-projects queries, keys, and values and performs attention in a compressed latent space. It targets the compute and KV-cache costs of long-context transformers.

    Engineers working on long-context models can assess an attention approach aimed at reducing both compute and cache costs.

  5. Continuous Thought Machines introduce neuron-level dynamics

    The paper presents the Continuous Thought Machine, a model that uses neuron-level processing and synchronization as core representations, drawing on neural dynamics in biological brains.

    Engineers exploring alternatives to standard neural-network architectures can examine its use of neuron timing and synchronization.

  6. MoE-CL for continual instruction tuning

    The paper proposes MoE-CL, a parameter-efficient mixture-of-experts framework for continual instruction tuning, combining task-specific and shared LoRA experts with a task-aware discriminator. It reports evaluations on MTL5 and Tencent3 and an A/B test at Tencent Video.

    The approach addresses catastrophic forgetting while adapting LLMs to changing tasks.

  7. Study compares in-context and in-weight learning

    The post describes a study in which a model trained on 12,000 small tasks combined new ideas. It reports a flexibility–memory trade-off and different effects from structured versus mixed lessons.

    The findings may inform how training curricula balance adaptation to new examples with retaining learned information.

  8. RLAD Trains LLMs to Generate Reasoning Hints

    The paper introduces a hint generator and a hint-conditioned solver, trained with reinforcement learning to produce and use short guidance for reasoning. The post reports improved math-task performance, including 44% higher AIME 2025 accuracy than long-chain RL.

    It explores whether reusable guidance can improve reasoning without relying on longer chains of thought.

  9. Slides Relate Kernel Regression to Attention

    A 10-slide summary examines how kernel regression is related to the attention mechanism.

    The connection can help engineers reason about attention through the lens of kernel methods.

  10. A three-stage account of grokking in modular addition

    The post describes a paper explaining grokking in modular addition as three training stages, and says generalization can emerge with about group-size times log(group-size) data. It highlights weight decay and moderate width as conditions for feature learning.

    The proposed stages and data threshold may help engineers reason about when models move from memorization to generalization.

  11. RLAD uses two-player RL to discover reasoning abstractions

    The post introduces RLAD, a two-player reinforcement learning framework for LLMs to discover natural-language hints that encode procedural knowledge for structured reasoning exploration.

    The framework explores a way to help LLMs structure reasoning through learned procedural hints.

  12. Dragon Hatchling Uses Local Neuron Rules for Language Modeling

    The post describes Dragon Hatchling, a language model built from simple neurons whose connection strengths store short-term memory through Hebbian learning. It also mentions BDH-GPU, a GPU-friendly version that behaves like an attention model with a running state.

    The architecture offers an alternative to standard attention and makes its memory and activations more inspectable.

  13. Dragon Hatchling uses locally interacting neuron particles

    Dragon Hatchling (BDH) is an LLM architecture based on a scale-free, biologically inspired network of locally interacting neuron particles. The post says it rivals GPT-2 performance and is designed for interpretability.

    Its architecture offers engineers an alternative design point focused on interpretability.

  14. Energy-Based Transformers refine next-token predictions with gradient steps

    Energy-Based Transformers score candidate next tokens by energy and iteratively lower that energy with gradient steps. In 44-million-parameter trials on RedPajama-Data-v2, they beat same-size vanilla transformers on three of four benchmarks.

    The approach offers an alternative to one-shot next-token prediction and reports benchmark results against vanilla transformers.

  15. A Unified View of Attention-MoE and FFN-MoE

    The post describes a paper that treats pre-mixed tokens and (V·Wo) as an FFN expert, unifying attention-MoE and FFN-MoE with shared experts. It claims lower perplexity at similar compute.

    The proposed framing may help engineers compare sparse attention and FFN designs.

  16. Gradient descent as attention over training examples

    The paper expresses gradient-trained linear layers as key-value memories containing training datapoints and initial weights. Predictions use unnormalized dot-product attention over the training experience.

    This view connects training dynamics to attention and may help engineers reason about what learned parameters store.

Build with AgentLog

List your MCP, skill or plugin

Reach the engineers who read these briefings.

Sponsor AgentLog

Footer, sidebar or featured slot for 30 days.

From US$ 60

See the slots

Send your own newsletter

CommsHarbor keeps contacts, consent and one-click unsubscribe together.

Free workspace

Open CommsHarbor