C++ has been the go-to for GPU programming for almost 20 years. Can Rust do the job, and how well?This book is all about getting hands-on with different toolchains that connect Rust to NVIDIA hardware. There's RustaCUDA for safe host-side control, the Rust-CUDA project for writing kernels in pure Rust, and NVIDIA's experimental cuda-oxide compiler with its typed launches and async execution graphs.
We're going to build one Cargo workspace that keeps on growing. It'll include device queries, launch planning, Rust-written kernels, memory optimization, parallel reductions and scans, multi-stream pipelines, matrix multiplication benchmarked against cuBLAS, a Monte Carlo option pricer validated against a closed formula, and a complete batched inference application measured against a Python baseline.
We'll check every result against a CPU reference, and the reports will give accurate numbers, including where libraries outperform hand-written kernels and where experimental toolchains are still a work in progress. Key LearningsLaunch, synchronize, and verify GPU kernels with ownership-managed device memory. Write real CUDA kernels using Rust-CUDA and cuda-oxide. Plan grids, blocks, and warps for 2D workloads.
Accelerate transfer speeds with pinned memory and coalesced access patterns. Build race-free thread cooperation using shared memory, barriers, and atomics. Overlap transfers with computation using streams, events, and async Rust pipelines. Optimize matrix multiplication and benchmark against cuBLAS ceiling. Wrap CUDA C library safely with handles, error enums, and Drop. Ship complete batched GPU inference application against Python baselines.
Diagnose performance with Nsight Systems, Nsight Compute, and compute-sanitizer. Table of ContentNew Beneficiary of GPU ComputingThinking in ThreadsCommanding GPUWriting GPU KernelsCleaner Kernels with cuda-oxideMastering GPU MemoryMaking Threads CooperateKeeping GPU BusyDelivering Real MathBorrowing NVIDIA's MuscleShipping Complete GPU ApplicationProving Performance
C++ has been the go-to for GPU programming for almost 20 years. Can Rust do the job, and how well?This book is all about getting hands-on with different toolchains that connect Rust to NVIDIA hardware. There's RustaCUDA for safe host-side control, the Rust-CUDA project for writing kernels in pure Rust, and NVIDIA's experimental cuda-oxide compiler with its typed launches and async execution graphs.
We're going to build one Cargo workspace that keeps on growing. It'll include device queries, launch planning, Rust-written kernels, memory optimization, parallel reductions and scans, multi-stream pipelines, matrix multiplication benchmarked against cuBLAS, a Monte Carlo option pricer validated against a closed formula, and a complete batched inference application measured against a Python baseline.
We'll check every result against a CPU reference, and the reports will give accurate numbers, including where libraries outperform hand-written kernels and where experimental toolchains are still a work in progress. Key LearningsLaunch, synchronize, and verify GPU kernels with ownership-managed device memory. Write real CUDA kernels using Rust-CUDA and cuda-oxide. Plan grids, blocks, and warps for 2D workloads.
Accelerate transfer speeds with pinned memory and coalesced access patterns. Build race-free thread cooperation using shared memory, barriers, and atomics. Overlap transfers with computation using streams, events, and async Rust pipelines. Optimize matrix multiplication and benchmark against cuBLAS ceiling. Wrap CUDA C library safely with handles, error enums, and Drop. Ship complete batched GPU inference application against Python baselines.
Diagnose performance with Nsight Systems, Nsight Compute, and compute-sanitizer. Table of ContentNew Beneficiary of GPU ComputingThinking in ThreadsCommanding GPUWriting GPU KernelsCleaner Kernels with cuda-oxideMastering GPU MemoryMaking Threads CooperateKeeping GPU BusyDelivering Real MathBorrowing NVIDIA's MuscleShipping Complete GPU ApplicationProving Performance