Attention and model architecture
Transformers, attention variants, state-space models and the ideas behind new architectures.
156 links, newest first.
- Attention and model architecturePost on X
Attention Sinks and Compression Valleys in LLMs
Experiments across models from 410M to 120B parameters report that extreme middle-layer activation norms at the beginning-of-sequence token coincide with compression valleys and attention sinks. The authors propose three phases of Transformer computation: broad mixing, limited mixing, and…
The proposed phase model may help engineers reason about how information mixing changes across Transformer layers.
- Attention and model architecturePost on X
Artificial Hippocampus Networks for Long-Context LLMs
A post says ByteDance released Artificial Hippocampus Networks (AHN), an architecture for long-context LLMs that continuously compresses information outside the context window.
Its approach to retaining out-of-window information may interest engineers exploring memory-efficient long-context architectures.
- Attention and model architecturePost on X
CCA Runs Attention in a Compressed Latent Space
The post describes Compressed Contextual Attention (CCA), which computes attention in a smaller latent space and applies RoPE there. It reports reduced KV-cache, parameter, and FLOP costs, plus faster prefill and backward passes in H100 tests.
CCA offers an attention design engineers can evaluate when balancing compute, memory, and model quality.
Compressed Convolutional Attention for Efficient Attention
The paper introduces Compressed Convolutional Attention (CCA), which down-projects queries, keys, and values and performs attention in a compressed latent space. It targets the compute and KV-cache costs of long-context transformers.
Engineers working on long-context models can assess an attention approach aimed at reducing both compute and cache costs.
Continuous Thought Machines introduce neuron-level dynamics
The paper presents the Continuous Thought Machine, a model that uses neuron-level processing and synchronization as core representations, drawing on neural dynamics in biological brains.
Engineers exploring alternatives to standard neural-network architectures can examine its use of neuron timing and synchronization.
MoE-CL for continual instruction tuning
The paper proposes MoE-CL, a parameter-efficient mixture-of-experts framework for continual instruction tuning, combining task-specific and shared LoRA experts with a task-aware discriminator. It reports evaluations on MTL5 and Tencent3 and an A/B test at Tencent Video.
The approach addresses catastrophic forgetting while adapting LLMs to changing tasks.
- Attention and model architecturePost on X
Study compares in-context and in-weight learning
The post describes a study in which a model trained on 12,000 small tasks combined new ideas. It reports a flexibility–memory trade-off and different effects from structured versus mixed lessons.
The findings may inform how training curricula balance adaptation to new examples with retaining learned information.
- Attention and model architecturePost on X
RLAD Trains LLMs to Generate Reasoning Hints
The paper introduces a hint generator and a hint-conditioned solver, trained with reinforcement learning to produce and use short guidance for reasoning. The post reports improved math-task performance, including 44% higher AIME 2025 accuracy than long-chain RL.
It explores whether reusable guidance can improve reasoning without relying on longer chains of thought.
- Attention and model architecturePost on X
Slides Relate Kernel Regression to Attention
A 10-slide summary examines how kernel regression is related to the attention mechanism.
The connection can help engineers reason about attention through the lens of kernel methods.
- Attention and model architecturePost on X
A three-stage account of grokking in modular addition
The post describes a paper explaining grokking in modular addition as three training stages, and says generalization can emerge with about group-size times log(group-size) data. It highlights weight decay and moderate width as conditions for feature learning.
The proposed stages and data threshold may help engineers reason about when models move from memorization to generalization.
- Attention and model architecturePost on X
RLAD uses two-player RL to discover reasoning abstractions
The post introduces RLAD, a two-player reinforcement learning framework for LLMs to discover natural-language hints that encode procedural knowledge for structured reasoning exploration.
The framework explores a way to help LLMs structure reasoning through learned procedural hints.
- Attention and model architecturePost on X
Dragon Hatchling Uses Local Neuron Rules for Language Modeling
The post describes Dragon Hatchling, a language model built from simple neurons whose connection strengths store short-term memory through Hebbian learning. It also mentions BDH-GPU, a GPU-friendly version that behaves like an attention model with a running state.
The architecture offers an alternative to standard attention and makes its memory and activations more inspectable.
- Attention and model architecturePost on X
Dragon Hatchling uses locally interacting neuron particles
Dragon Hatchling (BDH) is an LLM architecture based on a scale-free, biologically inspired network of locally interacting neuron particles. The post says it rivals GPT-2 performance and is designed for interpretability.
Its architecture offers engineers an alternative design point focused on interpretability.
Energy-Based Transformers refine next-token predictions with gradient steps
Energy-Based Transformers score candidate next tokens by energy and iteratively lower that energy with gradient steps. In 44-million-parameter trials on RedPajama-Data-v2, they beat same-size vanilla transformers on three of four benchmarks.
The approach offers an alternative to one-shot next-token prediction and reports benchmark results against vanilla transformers.
- Attention and model architecturePost on X
A Unified View of Attention-MoE and FFN-MoE
The post describes a paper that treats pre-mixed tokens and (V·Wo) as an FFN expert, unifying attention-MoE and FFN-MoE with shared experts. It claims lower perplexity at similar compute.
The proposed framing may help engineers compare sparse attention and FFN designs.
Gradient descent as attention over training examples
The paper expresses gradient-trained linear layers as key-value memories containing training datapoints and initial weights. Predictions use unnormalized dot-product attention over the training experience.
This view connects training dynamics to attention and may help engineers reason about what learned parameters store.
Build with AgentLog
Send your own newsletter
CommsHarbor keeps contacts, consent and one-click unsubscribe together.
Free workspace
Open CommsHarbor