Skip to content
EN

CUDA and GPU programming

Kernels, memory, profiling and the tooling that makes NVIDIA GPUs fast.

80 links, newest first.

Get the weekly briefing

The best new links of the topics you pick, summarized with the source. At most one email a week.

Topics: CUDA and GPU programming

Before the first issue we email you to confirm; leaving takes one click. Sent with CommsHarbor. Privacy

  1. Understanding Latency Hiding on GPUs

    A resource by Vasily Volkov of UC Berkeley about GPU latency hiding and architecture. The author says it covers performance concepts not explored in depth by *Programming Massively Parallel Processors*.

    Useful for engineers looking to deepen their understanding of GPU architecture and performance.

  2. Stanford course on parallel programming, GPUs, and CUDA

    The post shares a Stanford course with 19 lessons totaling 24 hours. Topics include GPU architecture and CUDA, performance optimization, and multicore processors and architectures.

    It offers a structured resource for learning GPU and CUDA programming and performance optimization.

  3. Tracing a CUDA Kernel from Compilation to Execution

    The article traces a vector-add kernel from nvcc through to the warps that execute it.

    It connects CUDA compilation with how GPU warps execute a kernel.

  4. Modern GPU Programming for ML Systems: Online Book

    A curated online book based on a mini-series taught at CMU's ML Systems course. Topics include data layout swizzling, 3D TMA, and Blackwell programming.

    It offers engineers course material on GPU programming techniques for ML systems.

  5. PufferLib 4.0 reports multi-GPU scaling results

    The post reports 6.6× scaling on 8 GPUs in a worst-case test with a 32× parameter policy, and peak throughput above 100M steps per second on 8 RTX 5090 GPUs.

    The reported figures offer a multi-GPU performance point of reference for GPU-based workloads.

  6. Is Parallel Programming Hard, and What Can You Do About It?

    A resource about the challenges of parallel programming and ways to address them.

    Parallel programming concepts can help engineers reason about concurrency in GPU kernels.

  7. Fault-Resilient GPU Multi-Process Service

    Researchers describe a GPU multi-process service combining driver-level isolation that confines reachable MMU faults to the faulting client with VMM-based active-standby recovery. They report zero runtime overhead for the isolation mechanism.

    The approach addresses fault containment and recovery in GPU multi-process workloads.

  8. Tile Scheduling in CuTeDSL

    A blog post introducing how StaticPersistentTileScheduler works in CuTeDSL and why it is useful when writing persistent kernels.

    Understanding tile scheduling can help engineers reason about how persistent kernels are organized.

  9. Georgia Tech series on high-performance computer architecture

    The post recommends a six-part Georgia Tech series for learning high-performance computer architecture.

    It may help engineers build background in the architecture behind fast GPU computing.

  10. CuBridge Uses LLMs to Reconstruct Attention Kernels

    CuBridge is an LLM-based framework for understanding and reconstructing high-performance attention kernels. The linked preview notes that supporting diverse, evolving attention variants is challenging.

    It addresses the challenge of supporting attention variants with efficient CUDA implementations.

  11. Video Explores GPU Architecture

    A video titled “How do Graphics Cards Work? Exploring GPU Architecture” offers an overview of graphics card internals.

    GPU architecture provides useful background for engineers working with GPU software.

  12. A CUDA/PTX matmul kernel beats cuBLAS on B200

    Paul Chan describes a pure CUDA/PTX matmul kernel on B200 that he says is 6% faster than cuBLAS at M=N=K=8192. The post links to a technical blog and a GitHub repository.

    The blog and code offer engineers a concrete B200 matmul optimization to examine.

  13. pyptx provides a Python DSL for NVIDIA PTX kernels

    pyptx is a Python DSL for writing NVIDIA PTX kernels, with Hopper and Blackwell support and JAX and PyTorch integration. It includes GEMM, grouped GEMM, RMSNorm, and SwiGLU examples.

    It offers a Python interface for writing PTX kernels and integrating them with JAX or PyTorch.

  14. GP-SIMD Processing-in-Memory

    The post links to an ACM paper identified as “GP-SIMD Processing-in-Memory.”

    It may interest engineers exploring GPU-style SIMD computing and processing-in-memory.

  15. Hugging Face’s Ultra-Scale Playbook on LLM Training

    The playbook covers data, expert, tensor, pipeline, and context parallelism for training LLMs on GPU clusters. It includes empirical examples from 4,000 scaling experiments across up to 512 GPUs.

    It connects parallelism choices to memory use, pipeline overhead, and cluster topology.

  16. Bindless access with 64-bit pointers

    The post says 64-bit pointers avoid fetching a buffer descriptor from a descriptor heap, describing this approach in CUDA and Metal and via Vulkan BDA in GLSL. It says DX12 has no pointer support.

    Useful context for comparing buffer access models across GPU APIs.

  17. A Primer on GPU Architecture and Computing

    An introductory article covering GPU architecture and computing.

    Useful background for engineers building or optimizing GPU software.

  18. Overlapping Data Transfers in CUDA C/C++

    An NVIDIA Technical Blog post by Mark Harris discusses how to overlap data transfers with computation in CUDA C/C++.

    Overlapping transfers and computation can help engineers make better use of host and GPU time.

  19. Comparing CPU and GPU Performance for ONNX Inference

    The article compares ONNX Runtime inference on an NVIDIA L4 GPU and an AMD EPYC 9965 CPU. The author says OCaml and libnuma memory placement made CPU inference faster than using scarce GPUs for TESSERA.

    Useful for engineers evaluating inference performance and memory placement across CPUs and GPUs.

  20. FP4 Training for Large-Scale MoE Models on Hopper

    The linked page discusses practical FP4 training for large-scale Mixture-of-Experts models on Hopper GPUs. It identifies activation memory and expert-parallel communication as training bottlenecks.

    Relevant to engineers working on GPU memory use and communication in large-scale MoE training.

Build with AgentLog

List your MCP, skill or plugin

Reach the engineers who read these briefings.

Sponsor AgentLog

Footer, sidebar or featured slot for 30 days.

From US$ 60

See the slots

Send your own newsletter

CommsHarbor keeps contacts, consent and one-click unsubscribe together.

Free workspace

Open CommsHarbor