CUDA and GPU programming
Kernels, memory, profiling and the tooling that makes NVIDIA GPUs fast.
80 links, newest first.
- CUDA and GPU programmingArticle
CUDA Agent uses agentic RL for CUDA kernel generation
The linked item presents CUDA Agent, a system for large-scale agentic reinforcement learning aimed at generating high-performance CUDA kernels. It notes that GPU kernel optimization requires specialized hardware expertise.
It may interest engineers exploring reinforcement learning approaches to CUDA kernel optimization.
- CUDA and GPU programmingArticle
Khronos Tutorial on Vulkan Subgroups
The tutorial explains Vulkan 1.1 subgroup functionality for sharing and manipulating data among parallel GPU tasks.
Subgroups provide a GPU parallelism mechanism engineers can compare with shared-memory approaches.
Using CUDA Warp-Level Primitives
An NVIDIA Technical Blog post about CUDA warp-level primitives. It describes how GPUs execute groups of threads called warps in SIMT fashion and how CUDA programs can take advantage of warp execution.
Understanding warp execution can help engineers write high-performance CUDA programs.
- CUDA and GPU programmingArticle
Vulkan Memory Barriers and Image Layouts Explained
A RasterGrid blog post explaining Vulkan memory barriers and image layouts.
Useful for engineers working with GPU memory synchronization in Vulkan.
- CUDA and GPU programmingArticle
Inside NVIDIA’s Blackwell GPU
The article examines NVIDIA’s Blackwell GPU and places it in the company’s tradition of building large GPUs.
GPU hardware context can help engineers understand the architecture behind their workloads.
- CUDA and GPU programmingPost on X
A CuPy/CUDA FlashAttention Implementation from Scratch
The project implements tiled FlashAttention with online softmax in CuPy and custom CUDA kernels, avoiding materialization of the NxN score matrix. The author reports mathematical verification but no speedup over cuBLAS.
It describes a hands-on implementation of FlashAttention and notes the performance gap with cuBLAS.
- CUDA and GPU programmingArticle
A comparison of inference-focused AI hardware
Turing Post compares the 2026 inference chip landscape, including NVIDIA Vera Rubin, MatX’s programmable LLM accelerator, and Taalas’ model-as-hardware approach, with a focus on cost per token.
Useful for engineers evaluating specialized accelerators and their trade-offs against GPU infrastructure.
A Step-by-Step Introduction to CUDA C++
NVIDIA's blog post introduces CUDA programming with a simple, step-by-step parallel programming example.
Useful for engineers looking for an accessible starting point for writing GPU code with CUDA C++.
- CUDA and GPU programmingArticle
A Worklog on Optimizing CUDA Matrix Multiplication
Simon Boehm documents the iterative optimization of a CUDA matrix multiplication kernel, aiming to understand how to approach cuBLAS-like performance rather than build a cuBLAS replacement.
The worklog offers a concrete view of kernel optimization decisions for a common GPU operation.
- CUDA and GPU programmingPost on X
Padding inputs to reuse variable-size CUDA graphs
The post recommends setting a maximum input size and padding inputs to the nearest interval, creating a family of graphs for different sequence lengths. It claims this avoids recompilation.
This is a concrete graph-capture strategy to consider for inference workloads with variable sequence lengths.
Efficient Matrix Transpose in CUDA C/C++
Mark Harris's NVIDIA blog post discusses performance gains achievable by using shared memory for matrix transpose in CUDA C/C++.
It explains how shared memory can improve the performance of a common GPU operation.
- CUDA and GPU programmingPost on X
CUDA Kernel Patterns: Reduction, Convolution, Histogram, and Softmax
The post lists parallel reduction, 1D convolution, histogram, and numerically stable softmax as core CUDA kernel patterns, mentioning techniques such as tiling, halo regions, data reuse, atomics, and multi-stage reduction.
These examples highlight common parallelization, memory-reuse, and contention challenges in CUDA kernels.
- CUDA and GPU programmingArticle
HPC++ Automatically Parallelizes C++ for CPUs and OpenCL GPUs
HPC++ is an LLVM-based framework that transforms sequential C++ programs into parallel implementations for multi-core CPUs and OpenCL-capable GPUs.
It may be relevant to engineers evaluating automatic parallelization across heterogeneous CPU–GPU systems.
- CUDA and GPU programmingRepository
Booth compiler targets multiple GPU and CPU architectures
Booth is an open-source compiler for CUDA, Triton, and HIP that targets multiple GPU and CPU architectures.
Engineers can explore a compiler project aimed at running these programming models across different architectures.
- CUDA and GPU programmingRepository
Booth is an open-source compiler for CUDA, Triton, and HIP
Booth targets multiple GPU and CPU architectures. The post describes BarraCUDA as compiling CUDA .cu files to AMD GFX11 machine code and ELF .hsaco binaries without an LLVM dependency.
It may help engineers explore compiling CUDA code for AMD GPUs and other architectures.
- CUDA and GPU programmingArticle
Inside VOLT: Designing an Open-Source GPU Compiler
An article about VOLT, an open-source GPU compiler. Its preview describes emerging open GPU architectures that define SIMT functionality.
Useful context for engineers exploring open-source GPU compiler design and architectures.
- CUDA and GPU programmingArticle
How GPUs Evolved from SIMD to SIMT
An ACM SIGGRAPH Blog post on how GPUs combined thread abstractions with SIMD hardware, becoming flexible processors for graphics, AI, and scientific computing.
Useful context for understanding the GPU execution model behind modern workloads.
- CUDA and GPU programmingArticle
CUDA for LLMs covers kernels, Flash Attention, and profiling
The Manning book introduces CUDA programming for NVIDIA GPUs, from first kernels to advanced LLM features such as Flash Attention. Its preview also mentions profiling with Nsight Compute.
Useful for engineers who want to learn GPU-level programming and profiling for LLM workloads.
CuTe Layouts Explain Matrix Shapes and Strides
The post recommends CUTLASS/CuTe’s Layout Algebra tutorials for understanding matrix shapes and strides. The linked NVIDIA documentation covers CuTe layouts.
Useful for engineers learning how CuTe represents tensor layouts in CUTLASS.
CUDA-L2 uses reinforcement learning to optimize HGEMM kernels
CUDA-L2 combines large language models and reinforcement learning to optimize half-precision General Matrix Multiply CUDA kernels, using execution speed as the reward. The paper evaluates kernels across 1,000 configurations and reports outperforming matmul baselines including cuBLAS.
The approach offers a way to automate tuning CUDA matrix multiplication kernels against execution-speed measurements.
Build with AgentLog
Send your own newsletter
CommsHarbor keeps contacts, consent and one-click unsubscribe together.
Free workspace
Open CommsHarbor







