CUDA and GPU programming
Kernels, memory, profiling and the tooling that makes NVIDIA GPUs fast.
80 links, newest first.
- CUDA and GPU programmingPost on X
Understanding Latency Hiding on GPUs
A resource by Vasily Volkov of UC Berkeley about GPU latency hiding and architecture. The author says it covers performance concepts not explored in depth by *Programming Massively Parallel Processors*.
Useful for engineers looking to deepen their understanding of GPU architecture and performance.
Stanford course on parallel programming, GPUs, and CUDA
The post shares a Stanford course with 19 lessons totaling 24 hours. Topics include GPU architecture and CUDA, performance optimization, and multicore processors and architectures.
It offers a structured resource for learning GPU and CUDA programming and performance optimization.
- CUDA and GPU programmingArticle
Tracing a CUDA Kernel from Compilation to Execution
The article traces a vector-add kernel from nvcc through to the warps that execute it.
It connects CUDA compilation with how GPU warps execute a kernel.
- CUDA and GPU programmingArticle
Modern GPU Programming for ML Systems: Online Book
A curated online book based on a mini-series taught at CMU's ML Systems course. Topics include data layout swizzling, 3D TMA, and Blackwell programming.
It offers engineers course material on GPU programming techniques for ML systems.
- CUDA and GPU programmingPost on X
PufferLib 4.0 reports multi-GPU scaling results
The post reports 6.6× scaling on 8 GPUs in a worst-case test with a 32× parameter policy, and peak throughput above 100M steps per second on 8 RTX 5090 GPUs.
The reported figures offer a multi-GPU performance point of reference for GPU-based workloads.
- CUDA and GPU programmingArticle
Is Parallel Programming Hard, and What Can You Do About It?
A resource about the challenges of parallel programming and ways to address them.
Parallel programming concepts can help engineers reason about concurrency in GPU kernels.
Fault-Resilient GPU Multi-Process Service
Researchers describe a GPU multi-process service combining driver-level isolation that confines reachable MMU faults to the faulting client with VMM-based active-standby recovery. They report zero runtime overhead for the isolation mechanism.
The approach addresses fault containment and recovery in GPU multi-process workloads.
- CUDA and GPU programmingArticle
Tile Scheduling in CuTeDSL
A blog post introducing how StaticPersistentTileScheduler works in CuTeDSL and why it is useful when writing persistent kernels.
Understanding tile scheduling can help engineers reason about how persistent kernels are organized.
- CUDA and GPU programmingPost on X
Georgia Tech series on high-performance computer architecture
The post recommends a six-part Georgia Tech series for learning high-performance computer architecture.
It may help engineers build background in the architecture behind fast GPU computing.
- CUDA and GPU programmingArticle
CuBridge Uses LLMs to Reconstruct Attention Kernels
CuBridge is an LLM-based framework for understanding and reconstructing high-performance attention kernels. The linked preview notes that supporting diverse, evolving attention variants is challenging.
It addresses the challenge of supporting attention variants with efficient CUDA implementations.
Video Explores GPU Architecture
A video titled “How do Graphics Cards Work? Exploring GPU Architecture” offers an overview of graphics card internals.
GPU architecture provides useful background for engineers working with GPU software.
- CUDA and GPU programmingArticle
A CUDA/PTX matmul kernel beats cuBLAS on B200
Paul Chan describes a pure CUDA/PTX matmul kernel on B200 that he says is 6% faster than cuBLAS at M=N=K=8192. The post links to a technical blog and a GitHub repository.
The blog and code offer engineers a concrete B200 matmul optimization to examine.
- CUDA and GPU programmingRepository
pyptx provides a Python DSL for NVIDIA PTX kernels
pyptx is a Python DSL for writing NVIDIA PTX kernels, with Hopper and Blackwell support and JAX and PyTorch integration. It includes GEMM, grouped GEMM, RMSNorm, and SwiGLU examples.
It offers a Python interface for writing PTX kernels and integrating them with JAX or PyTorch.
GP-SIMD Processing-in-Memory
The post links to an ACM paper identified as “GP-SIMD Processing-in-Memory.”
It may interest engineers exploring GPU-style SIMD computing and processing-in-memory.
- CUDA and GPU programmingPost on X
Hugging Face’s Ultra-Scale Playbook on LLM Training
The playbook covers data, expert, tensor, pipeline, and context parallelism for training LLMs on GPU clusters. It includes empirical examples from 4,000 scaling experiments across up to 512 GPUs.
It connects parallelism choices to memory use, pipeline overhead, and cluster topology.
- CUDA and GPU programmingPost on X
Bindless access with 64-bit pointers
The post says 64-bit pointers avoid fetching a buffer descriptor from a descriptor heap, describing this approach in CUDA and Metal and via Vulkan BDA in GLSL. It says DX12 has no pointer support.
Useful context for comparing buffer access models across GPU APIs.
- CUDA and GPU programmingArticle
A Primer on GPU Architecture and Computing
An introductory article covering GPU architecture and computing.
Useful background for engineers building or optimizing GPU software.
Overlapping Data Transfers in CUDA C/C++
An NVIDIA Technical Blog post by Mark Harris discusses how to overlap data transfers with computation in CUDA C/C++.
Overlapping transfers and computation can help engineers make better use of host and GPU time.
- CUDA and GPU programmingArticle
Comparing CPU and GPU Performance for ONNX Inference
The article compares ONNX Runtime inference on an NVIDIA L4 GPU and an AMD EPYC 9965 CPU. The author says OCaml and libnuma memory placement made CPU inference faster than using scarce GPUs for TESSERA.
Useful for engineers evaluating inference performance and memory placement across CPUs and GPUs.
- CUDA and GPU programmingArticle
FP4 Training for Large-Scale MoE Models on Hopper
The linked page discusses practical FP4 training for large-scale Mixture-of-Experts models on Hopper GPUs. It identifies activation memory and expert-parallel communication as training bottlenecks.
Relevant to engineers working on GPU memory use and communication in large-scale MoE training.
Build with AgentLog
Send your own newsletter
CommsHarbor keeps contacts, consent and one-click unsubscribe together.
Free workspace
Open CommsHarbor





