Skip to content
EN

CUDA and GPU programming

Kernels, memory, profiling and the tooling that makes NVIDIA GPUs fast.

80 links, newest first.

Get the weekly briefing

The best new links of the topics you pick, summarized with the source. At most one email a week.

Topics: CUDA and GPU programming

Before the first issue we email you to confirm; leaving takes one click. Sent with CommsHarbor. Privacy

  1. CUDA Agent uses agentic RL for CUDA kernel generation

    The linked item presents CUDA Agent, a system for large-scale agentic reinforcement learning aimed at generating high-performance CUDA kernels. It notes that GPU kernel optimization requires specialized hardware expertise.

    It may interest engineers exploring reinforcement learning approaches to CUDA kernel optimization.

  2. Khronos Tutorial on Vulkan Subgroups

    The tutorial explains Vulkan 1.1 subgroup functionality for sharing and manipulating data among parallel GPU tasks.

    Subgroups provide a GPU parallelism mechanism engineers can compare with shared-memory approaches.

  3. Using CUDA Warp-Level Primitives

    An NVIDIA Technical Blog post about CUDA warp-level primitives. It describes how GPUs execute groups of threads called warps in SIMT fashion and how CUDA programs can take advantage of warp execution.

    Understanding warp execution can help engineers write high-performance CUDA programs.

  4. Vulkan Memory Barriers and Image Layouts Explained

    A RasterGrid blog post explaining Vulkan memory barriers and image layouts.

    Useful for engineers working with GPU memory synchronization in Vulkan.

  5. Inside NVIDIA’s Blackwell GPU

    The article examines NVIDIA’s Blackwell GPU and places it in the company’s tradition of building large GPUs.

    GPU hardware context can help engineers understand the architecture behind their workloads.

  6. A CuPy/CUDA FlashAttention Implementation from Scratch

    The project implements tiled FlashAttention with online softmax in CuPy and custom CUDA kernels, avoiding materialization of the NxN score matrix. The author reports mathematical verification but no speedup over cuBLAS.

    It describes a hands-on implementation of FlashAttention and notes the performance gap with cuBLAS.

  7. A comparison of inference-focused AI hardware

    Turing Post compares the 2026 inference chip landscape, including NVIDIA Vera Rubin, MatX’s programmable LLM accelerator, and Taalas’ model-as-hardware approach, with a focus on cost per token.

    Useful for engineers evaluating specialized accelerators and their trade-offs against GPU infrastructure.

  8. A Step-by-Step Introduction to CUDA C++

    NVIDIA's blog post introduces CUDA programming with a simple, step-by-step parallel programming example.

    Useful for engineers looking for an accessible starting point for writing GPU code with CUDA C++.

  9. A Worklog on Optimizing CUDA Matrix Multiplication

    Simon Boehm documents the iterative optimization of a CUDA matrix multiplication kernel, aiming to understand how to approach cuBLAS-like performance rather than build a cuBLAS replacement.

    The worklog offers a concrete view of kernel optimization decisions for a common GPU operation.

  10. Padding inputs to reuse variable-size CUDA graphs

    The post recommends setting a maximum input size and padding inputs to the nearest interval, creating a family of graphs for different sequence lengths. It claims this avoids recompilation.

    This is a concrete graph-capture strategy to consider for inference workloads with variable sequence lengths.

  11. Efficient Matrix Transpose in CUDA C/C++

    Mark Harris's NVIDIA blog post discusses performance gains achievable by using shared memory for matrix transpose in CUDA C/C++.

    It explains how shared memory can improve the performance of a common GPU operation.

  12. CUDA Kernel Patterns: Reduction, Convolution, Histogram, and Softmax

    The post lists parallel reduction, 1D convolution, histogram, and numerically stable softmax as core CUDA kernel patterns, mentioning techniques such as tiling, halo regions, data reuse, atomics, and multi-stage reduction.

    These examples highlight common parallelization, memory-reuse, and contention challenges in CUDA kernels.

  13. HPC++ Automatically Parallelizes C++ for CPUs and OpenCL GPUs

    HPC++ is an LLVM-based framework that transforms sequential C++ programs into parallel implementations for multi-core CPUs and OpenCL-capable GPUs.

    It may be relevant to engineers evaluating automatic parallelization across heterogeneous CPU–GPU systems.

  14. Booth compiler targets multiple GPU and CPU architectures

    Booth is an open-source compiler for CUDA, Triton, and HIP that targets multiple GPU and CPU architectures.

    Engineers can explore a compiler project aimed at running these programming models across different architectures.

  15. Booth is an open-source compiler for CUDA, Triton, and HIP

    Booth targets multiple GPU and CPU architectures. The post describes BarraCUDA as compiling CUDA .cu files to AMD GFX11 machine code and ELF .hsaco binaries without an LLVM dependency.

    It may help engineers explore compiling CUDA code for AMD GPUs and other architectures.

  16. Inside VOLT: Designing an Open-Source GPU Compiler

    An article about VOLT, an open-source GPU compiler. Its preview describes emerging open GPU architectures that define SIMT functionality.

    Useful context for engineers exploring open-source GPU compiler design and architectures.

  17. How GPUs Evolved from SIMD to SIMT

    An ACM SIGGRAPH Blog post on how GPUs combined thread abstractions with SIMD hardware, becoming flexible processors for graphics, AI, and scientific computing.

    Useful context for understanding the GPU execution model behind modern workloads.

  18. CUDA for LLMs covers kernels, Flash Attention, and profiling

    The Manning book introduces CUDA programming for NVIDIA GPUs, from first kernels to advanced LLM features such as Flash Attention. Its preview also mentions profiling with Nsight Compute.

    Useful for engineers who want to learn GPU-level programming and profiling for LLM workloads.

  19. CuTe Layouts Explain Matrix Shapes and Strides

    The post recommends CUTLASS/CuTe’s Layout Algebra tutorials for understanding matrix shapes and strides. The linked NVIDIA documentation covers CuTe layouts.

    Useful for engineers learning how CuTe represents tensor layouts in CUTLASS.

  20. CUDA-L2 uses reinforcement learning to optimize HGEMM kernels

    CUDA-L2 combines large language models and reinforcement learning to optimize half-precision General Matrix Multiply CUDA kernels, using execution speed as the reward. The paper evaluates kernels across 1,000 configurations and reports outperforming matmul baselines including cuBLAS.

    The approach offers a way to automate tuning CUDA matrix multiplication kernels against execution-speed measurements.

Build with AgentLog

List your MCP, skill or plugin

Reach the engineers who read these briefings.

Sponsor AgentLog

Footer, sidebar or featured slot for 30 days.

From US$ 60

See the slots

Send your own newsletter

CommsHarbor keeps contacts, consent and one-click unsubscribe together.

Free workspace

Open CommsHarbor