Skip to content
EN

LLM inference

Serving, quantization, batching, speculative decoding and the cost of every token.

85 links, newest first.

Get the weekly briefing

The best new links of the topics you pick, summarized with the source. At most one email a week.

Topics: LLM inference

Before the first issue we email you to confirm; leaving takes one click. Sent with CommsHarbor. Privacy

  1. Dual-Cache enables cross-layer communication between LLMs

    The paper presents Dual-Cache, a latent protocol for transferring information between heterogeneous language models by translating a Sharer’s KV cache into a Receiver’s. The post says it adds joint cross-layer mixing.

    Engineers building multi-agent LLM systems can assess a way to share context without exchanging text.

  2. SVD-LLM Uses Truncation-Aware SVD for LLM Compression

    The paper proposes an SVD-based method for LLM compression that addresses compression loss from truncating smaller singular values and updates compressed weights after truncation.

    It describes an alternative to quantization for reducing model size, relevant to practical LLM deployment.

  3. LLM inferencePost on X

    How vLLM Schedules Requests and Allocates KV Cache

    An article traces a vLLM request from queueing through prefill and decoding. It explains how iteration-level scheduling and PagedAttention allocate KV cache blocks as sequences grow.

    Useful for understanding how vLLM manages batching and KV cache memory during inference.

  4. LLM inferenceRepository

    NVIDIA Model Optimizer combines model compression techniques

    NVIDIA Model Optimizer is a library for techniques including quantization, distillation, pruning, neural architecture search, and speculative decoding. The post says its exported checkpoints can be used with vLLM, SGLang, TensorRT-LLM, and TensorRT.

    Engineers can evaluate a single toolkit for reducing model size and preparing checkpoints for inference runtimes.

  5. PagedAttention manages KV cache memory for LLM serving

    The paper introduces PagedAttention, an attention algorithm inspired by virtual memory, to address KV cache fragmentation and redundant duplication that can limit batch size in LLM serving.

    It describes a memory-management approach that can help engineers understand constraints on serving batch size.

  6. Test-time training layers for sequence modeling

    The paper presents a framework for sequence modeling layers with linear complexity and expressive hidden states. Its key idea is to make the hidden state a machine learning model and update it through self-supervised learning.

    It explores an alternative to conventional attention for long-context sequence modeling.

  7. LLM inferenceRepository

    Zlaya: CPU-Only Inference Engine for Laya in Zig

    Zlaya is a CPU-only inference engine for Laya, implemented in Zig. Its GitHub repository describes it as fast and compact.

    Engineers exploring CPU-based Laya inference can inspect its Zig implementation.

  8. LLM inferencePost on X

    BITCOS Compresses Ternary LLMs to 1.485 Bits per Weight

    The post describes Intel’s BITCOS method, which uses zero weights in ternary LLMs to store their exact weights at 1.485 bits per weight. It reports up to 18% higher CPU and 27% higher GPU decode throughput.

    The reported compression and decode gains could matter for serving ternary LLMs on memory- and compute-constrained hardware.

  9. Speculative Decoding Speeds Up Autoregressive Inference

    The paper introduces speculative decoding, which uses an efficient model to propose several tokens and a larger model to verify them in parallel. It aims to speed up sampling without changing the outputs.

    It describes a method for reducing serial decoding steps while preserving the model’s output distribution.

  10. Crusoe reports MLPerf inference results on AMD MI355X

    Crusoe describes MLPerf Inference v6.1 results using 512 AMD Instinct MI355X GPUs, reporting 5.75 million tokens per second on gpt-oss-120b and linear scaling without InfiniBand.

    The results offer a reference point for evaluating inference-cluster throughput and scaling.

  11. Switch Attention dynamically combines full and windowed attention

    The paper introduces Switch Attention, a dynamic, fine-grained hybrid transformer approach. It addresses the quadratic cost of full attention and the narrower receptive fields of sliding-window attention.

    Engineers working on long-context inference can assess a dynamic alternative to fixed attention patterns.

  12. LLM inferenceRepository

    Open book covers AI infrastructure and system design

    An open book with 12 chapters on AI infrastructure, from model architectures and accelerators to inference optimization, distributed inference, training systems, and resource scheduling. It includes the full manuscript, a PDF, tools, and experiments.

    Useful as a broad reference for engineers working on LLM inference and training systems.

  13. LLM inferencePost on X

    Post describes DeepSeek’s claimed KV cache reduction

    The post says DeepSeek reduced KV cache use from 389,000 bytes to 890 bytes per token over three years, and says its linked article explains the approach with diagrams.

    KV cache size can affect memory use and the cost of serving long-context agents.

  14. LLM inferencePost on X

    DeepSeek-V4.1-Flash Runs Across GPU, RAM, and NVMe

    The post reports running DeepSeek-V4.1-Flash on one RTX 5090 with 125.7 GiB of RAM, while a 502GB GGUF resides partly on NVMe. The author reports 5.12 tok/s on new content and up to 21.27 tok/s when data is resident.

    The reported setup and throughput offer a concrete example of disk-backed model serving across GPU, RAM, and NVMe.

  15. LLM inferenceRepository

    Colibri streams MoE experts from disk in a C inference engine

    Colibri is a pure C engine with no dependencies that streams MoE experts from disk to run models on local hardware.

    Engineers can evaluate an approach to running MoE models with limited hardware memory.

  16. LLM inferenceRepository

    DeepSeek releases repositories for deploying V4.1 Flash

    DeepSeek says it has open-sourced repositories to make deploying V4.1 Flash and later open models easier. The linked repositories are deepseek-recipe, DeepSelect, and DeepJIT.

    The repositories may offer engineers resources for deploying DeepSeek models.

  17. Hy4 Preview Uses Sherry Quantization at 1.25 Bits per Weight

    Tencent’s post says Sherry quantization reduces the Hy4 preview from 1.5 TB to 214 GB, using 1.25 bits per weight. It also describes running the model across GPUs in multiple machines.

    The model and quantized GGUF are relevant to engineers exploring lower-memory inference and multi-machine GPU setups.

  18. LLM inferencePost on X

    KV Cache: Queries Are Not Retained

    The post says queries are used once, while keys and values are the parts retained in a KV cache.

    It gives a concise distinction relevant to understanding KV-cache behavior during inference.

  19. LLM inferencePost on X

    Post describes calibration-guided low-bit compression of Hy4

    The post says Tencent compressed the Hy4 preview from 1.5 TB to 200 GB by using calibration data to set layer-specific bit widths, with some layers at 1.31 bits. It reports MCP Atlas at 83.2 and SWE-Bench Multi at 81.3, down from 82.9.

    Layer-specific quantization can help engineers weigh model size reductions against benchmark changes.

  20. LLM inferenceRepository

    Colibri streams MoE model experts from disk in pure C

    Colibri is a pure C project for running seven open MoE model families, from 744B to 2.8T parameters, with experts streamed from disk. The post says it runs on hardware users own.

    Engineers can inspect an approach to serving large MoE models with disk-streamed experts.

Build with AgentLog

List your MCP, skill or plugin

Reach the engineers who read these briefings.

Sponsor AgentLog

Footer, sidebar or featured slot for 30 days.

From US$ 60

See the slots

Send your own newsletter

CommsHarbor keeps contacts, consent and one-click unsubscribe together.

Free workspace

Open CommsHarbor