Skip to content
EN

LLM inference

Serving, quantization, batching, speculative decoding and the cost of every token.

85 links, newest first.

Get the weekly briefing

The best new links of the topics you pick, summarized with the source. At most one email a week.

Topics: LLM inference

Before the first issue we email you to confirm; leaving takes one click. Sent with CommsHarbor. Privacy

  1. LLM inferenceRepository

    Community LLM serving recipes for RTX 3090, 4090, and 5090

    The GitHub repository provides model-agnostic community recipes for serving LLMs on RTX 3090, 4090, and 5090 GPUs, with support for vLLM, llama.cpp, and ik_llama.

    Engineers can compare serving setups across GPU generations and inference engines.

  2. A guide to 13 open-source foundation model deployment tools

    Turing Post’s guide covers 13 open-source tools for deploying, serving, and running foundation models, from local LLMs to high-throughput production inference.

    It helps engineers compare deployment tools across local and production inference use cases.

  3. LLM inferenceRepository

    llama.cpp Adds Multi-Token Prediction Support

    A llama.cpp pull request adds support for MTP heads, which the post says can predict multiple tokens per pass. The author reports Qwen3.6-27B reached 65 tok/s, up from 38 tok/s, on an RTX 3090 with MTP enabled.

    Engineers serving supported models can evaluate MTP as a way to increase inference throughput.

  4. LLM inferencePost on X

    MTP inference benchmarks on three GTX 1080 Ti GPUs

    The post reports llama.cpp MTP throughput benchmarks for Qwen 3.6 models on three GTX 1080 Ti GPUs, with Q4_0 K/V cache and draft-MTP flags. It lists context sizes and token rates.

    The reported setup and flags may help engineers compare inference performance on older GPUs.

  5. Fluxion: Hybrid Sparse Attention for Long-Context Inference

    The paper presents Fluxion, a hybrid sparse-attention system for long-context inference that uses CPU-GPU parallelism.

    Engineers working on long-context inference may find its CPU-GPU parallelism approach relevant.

  6. LLM inferencePost on X

    RTX 3090 LLM efficiency tests at different power limits

    The author reports tests of Qwen3.6 27B on one RTX 3090 at concurrencies of 1, 4, 8, and 16. They say 225W had the best efficiency, while 250W provided higher throughput.

    Useful measurements for engineers tuning power limits and concurrency on local LLM inference.

  7. LLM inferencePost on X

    Luce PFlash Reports 2.89× TTFT Speedup Over Ollama at 64K

    The post reports a benchmark in which Luce PFlash achieved 2.89× faster time to first token than Ollama at a 64K context.

    The result may be relevant when comparing inference serving performance at long context lengths.

  8. How KV cache avoids recomputing processed tokens

    The article explains that an LLM's KV cache stores the Key and Value of tokens it has already processed, so they do not need to be computed again for each new token.

    Understanding KV cache helps engineers reason about memory use and repeated computation during LLM inference.

  9. LLM inferenceRepository

    Lucebox: speculative inference server for consumer GPUs

    The post claims Qwen3.6-27B reaches 120–200 tokens/s on one RTX 3090 and links to Lucebox, a speculative inference server for heterogeneous hardware and consumer GPUs.

    The repository may help engineers explore speculative inference on consumer GPU hardware.

  10. LLM inferencePost on X

    RTX 3090 LLM tests compare throughput and power efficiency

    An engineer tested 8 local LLMs on a single RTX 3090 at power limits from 100W to 450W. Average throughput was 90.4 tok/s at 225W and 107.1 tok/s at 450W, while efficiency fell from 0.4167 to 0.2731 tok/s/W.

    The results show the throughput and power-efficiency tradeoff when choosing a GPU power limit for local inference.

  11. LLM inferenceRepository

    dflash-mlx v0.1.5 adds runtime and serving features

    The release adds a unified CLI for serving, generation, benchmarking, diagnostics, profiling, and model listing. It also introduces RAM and SSD snapshot caching, prefix matching, and experimental long-context support.

    Engineers serving models on Apple Silicon can evaluate its runtime, caching, and diagnostics features.

  12. LLM inferenceRepository

    Luce PFlash adds prefix caching and cold-start tuning

    Luce PFlash adds prefix caching, cold-start tuning, and a CUDA VMM fix. The author reports about 10× faster warm performance and 2.5× faster cold performance, with block sparse attention autotune, for Qwen3.6 27B.

    The linked repository is an LLM speculative inference server for heterogeneous hardware and consumer GPUs.

  13. LLM inferenceRepository

    Lucebox explores speculative prefill for faster LLM inference

    Lucebox is an LLM speculative inference server for heterogeneous hardware and consumer GPUs. The post claims speculative prefill speeds up Qwen3.6 27B time to first token by up to 10×.

    Engineers can inspect a server implementation focused on speculative inference across different hardware.

  14. LLM inferencePost on X

    Post claims DeepSeek-V4-Flash could fit on two RTX Pro 6000s

    The author says DeepSeek-V4-Flash’s experts are already 4-bit and claims the model would fit on two RTX Pro 6000s. They say vLLM or SGLang still needs SM120 support for it.

    It highlights how quantization and serving-stack support may affect the hardware needed for inference.

  15. LLM inferencePost on X

    DFlash and DDTree run Qwen3.6-27B on an RTX 3090

    The post reports 73 tok/s for Qwen3.6-27B on one RTX 3090 using a DFlash and DDTree speculative decoding stack. It says the stack loads the model because its architecture string and layer and head dimensions match Qwen3.5, but notes lower throughput than on Qwen3.5.

    It highlights how architecture compatibility can enable speculative decoding on consumer GPUs before dedicated upstream support arrives.

  16. Inference Engineering covers production AI inference systems

    Philip Kiely's book covers the hardware, software, techniques, and infrastructure required to run AI models in production. The post says it focuses mostly on LLMs, with some coverage of diffusion.

    A broad overview can help engineers build foundational knowledge of production inference systems.

  17. LLM inferenceRepository

    Training Neural Networks on Apple’s Neural Engine

    The ANE repository describes training neural networks using reverse-engineered private APIs for Apple’s Neural Engine.

    Engineers exploring neural-network training and inference on Apple hardware can examine the implementation.

  18. Fast LLM Inference From Scratch

    Andrew Chan’s article is titled “Fast LLM Inference From Scratch.”

    It may offer engineers a practical entry point to implementing LLM inference.

  19. OpenAnonymity Proposes Unlinkable Inference for AI

    The post describes a layer intended to prevent AI providers from linking inference calls to individual users. It says the approach is built into open-source infrastructure and a chat app.

    Engineers evaluating hosted LLMs can assess an approach to separating user identity from inference requests.

  20. LLM inferencePost on X

    Technical tutorial on dLLM inference and training

    A roughly 22-minute tutorial on dLLMs covers self-distillation to reduce diffusion steps, curriculum learning, an LLM verifier, and KV cache.

    It outlines training and inference techniques relevant to engineers exploring dLLMs.

Build with AgentLog

List your MCP, skill or plugin

Reach the engineers who read these briefings.

Sponsor AgentLog

Footer, sidebar or featured slot for 30 days.

From US$ 60

See the slots

Send your own newsletter

CommsHarbor keeps contacts, consent and one-click unsubscribe together.

Free workspace

Open CommsHarbor