Skip to content
EN

LLM inference

Serving, quantization, batching, speculative decoding and the cost of every token.

85 links, newest first.

Get the weekly briefing

The best new links of the topics you pick, summarized with the source. At most one email a week.

Topics: LLM inference

Before the first issue we email you to confirm; leaving takes one click. Sent with CommsHarbor. Privacy

  1. LLM inferenceRepository

    ntransformer Runs Llama 70B on an RTX 3090

    ntransformer is a C++/CUDA LLM inference engine whose GitHub page says it can run Llama 70B on an RTX 3090. The post describes using NVMe-to-GPU transfer while bypassing the CPU.

    Engineers can examine an inference engine targeting large-model execution on a single consumer GPU.

  2. LLM inferenceRepository

    llmfit identifies models that fit local hardware

    llmfit is a GitHub tool for finding models and providers that run on your hardware. The post says it probes hardware, selects quantization for available RAM, handles MoE expert offloading, and estimates tokens per second.

    It can help engineers assess local model options and expected inference speed before downloading weights.

  3. LLM inferencePost on X

    hf-mem estimates model memory requirements

    A post describes hf-mem as a tool for estimating memory requirements for new models and gives an example command using Qwen3.5-397B-A17B with experimental mode and an FP8 KV cache.

    It offers a command-line approach to estimating memory needs as model architectures change.

  4. Latency Optimization for Qwen3 on AMD MI300X

    An LMSYS blog post about latency optimization for Qwen3 and Qwen3-VL on AMD MI300X GPUs.

    It may offer relevant inference optimization techniques for engineers serving Qwen models on AMD hardware.

  5. Unsloth Guide to Tool Calling with Local LLMs

    Unsloth’s guide covers function calling with open models, with examples for story writing, Python execution, terminal calls and math.

    Useful for engineers integrating tool calling into applications that run local open models.

  6. LLM inferencePost on X

    LMCache reuses KV states across storage tiers

    LMCache is an open-source extension for LLM serving that manages KV cache across GPUs, CPUs, and local disks. It can reuse repeated text fragments, not only prefixes.

    Reusing KV states can reduce prefill work and GPU memory use in LLM serving.

  7. LLM inferenceRepository

    Tencent Releases HPC-Ops LLM Inference Operator Library

    HPC-Ops is an open-source LLM inference operator library with FusedMoE, GroupGEMM, attention, and multi-node communication support. Tencent reports production throughput gains and kernel speedups against named alternatives.

    Engineers can inspect its operators and benchmarks when optimizing inference on GPUs.

  8. LLM inferencePost on X

    CLI estimates model-loading VRAM requirements

    The `do-i-have-the-vram` tool estimates how much VRAM is needed to load a model without loading it. It can be installed with `pip install do-i-have-the-vram`.

    It can help engineers check whether a model fits on available GPU memory before attempting to load it.

  9. LoPA uses lookahead token ordering to speed up dLLM inference

    LoPA is a training-free, plug-and-play decoding algorithm for diffusion LLMs. It uses lookahead to identify token filling orders, addressing the limited parallelism of confidence-driven decoding.

    Token filling order can affect parallelism, making LoPA relevant to engineers optimizing diffusion LLM inference.

  10. GLM-4.7 185B W4A16 weights on Hugging Face

    The linked Hugging Face repository is for GLM-4.7 185B W4A16 weights. The post says the weights are 92 GB and that the author ran 100 million tokens locally.

    The quantized weights may be relevant to engineers evaluating local inference for a large model.

  11. Transformers optimizations used by OpenAI gpt-oss

    A Hugging Face blog post lists techniques from OpenAI’s gpt-oss model for use with Transformers, including MXFP4 quantization, tensor and expert parallelism, and dynamic sliding windows.

    These techniques are relevant to engineers optimizing LLM inference with Transformers.

  12. LLM inferencePost on X

    DEER Uses Diffusion Drafts for LLM Inference

    DEER drafts tokens with diffusion models and verifies them with autoregressive models. The post claims up to 5.54× faster inference with lossless acceleration.

    The approach may be relevant to engineers evaluating speculative decoding alternatives for inference acceleration.

  13. LLM inferencePost on X

    SWAA for Efficient Long-Context LLM Inference

    The post describes Sliding Window Attention Adaptation (SWAA), a set of training-free recipes for adapting full-attention LLMs to linear scaling while recovering performance.

    Engineers evaluating long-context inference can consider a training-free approach to reducing attention costs.

  14. LLM inferencePost on X

    Perplexity Builds MoE Inference Kernels for AWS EFA

    Perplexity describes expert-parallel kernels for serving large MoE models across AWS GPUs using EFA. GPU dispatch and combine pack tokens into RDMA writes, while a host proxy thread coordinates transfers alongside grouped GEMM compute.

    The design offers an approach to multi-node MoE serving when EFA lacks GPUDirect Async.

  15. Router-R1 Uses Reinforcement Learning to Coordinate Multiple LLMs

    Router-R1 frames multi-LLM routing as a sequential decision process, alternating between reasoning and invoking models. Its reward combines format, outcome, and cost terms; the post reports results across seven QA benchmarks.

    Engineers can examine an RL-based approach to balancing multi-model performance and cost.

  16. InfLLM-V2 Switches Between Dense and Sparse Attention

    InfLLM-V2 is a dense-sparse switchable attention system for adapting from short to long sequences. Its paper discusses long-sequence processing bottlenecks and limitations of existing trainable sparse attention methods.

    Engineers can assess an attention approach designed to support long contexts while addressing standard Transformer bottlenecks.

  17. LLM inferencePost on X

    PHLoRA extracts LoRA adapters from fine-tuned checkpoints

    PHLoRA uses the base and fine-tuned checkpoints to extract a low-rank adapter via singular value decomposition, without training data or gradients. The post reports accuracy within about 1% of full-rank models at ranks 32 or 64.

    Engineers can reduce adapter loading and serving costs by converting existing full-rank fine-tunes.

  18. Sipeed Maix4-HAT brings an AX650N NPU to Raspberry Pi

    The Raspberry Pi 5 AI HAT uses an AXera AX650N NPU, rated for up to 72 TOPS at INT4 or 18 TOPS at INT8. It includes 8 GB LPDDR4x RAM and accelerates Transformer-based models for edge applications.

    It offers engineers a compact platform to evaluate quantized Transformer inference at the edge.

  19. LLM inferenceRepository

    Mirage compiles LLMs into persistent megakernels

    Mirage Persistent Kernel is a compiler that transforms LLMs into optimized megakernels. The post claims this reduces latency by 1.2–6.7×.

    Engineers working on LLM inference can evaluate a compiler approach to reducing latency through megakernels.

  20. LLM inferencePost on X

    AirLLM runs large models with layer-wise inference

    AirLLM uses layer-wise inference, loading one Transformer layer at a time from disk so the full model need not stay in GPU memory. The post also describes block-wise quantization and compression support.

    Engineers evaluating inference on memory-constrained GPUs may find its layer-loading approach relevant.

Build with AgentLog

List your MCP, skill or plugin

Reach the engineers who read these briefings.

Sponsor AgentLog

Footer, sidebar or featured slot for 30 days.

From US$ 60

See the slots

Send your own newsletter

CommsHarbor keeps contacts, consent and one-click unsubscribe together.

Free workspace

Open CommsHarbor