CUDA is NVIDIA's programming platform for writing software that runs directly on its GPUs, and it's the foundation nearly every AI training and inference framework is built on top of.
A GPU can't run arbitrary code the way a CPU can; it needs code written specifically for its parallel architecture. CUDA is the toolkit and language extensions that let developers write that code, and it's what libraries like PyTorch use underneath to actually make an NVIDIA GPU do the math a neural network needs.
CUDA's decade-plus head start and deep integration into every major AI framework is a large part of why NVIDIA GPUs dominate AI infrastructure; competing hardware has to either support CUDA or convince the entire software ecosystem to retarget itself.
Developers write kernels, small functions meant to run on thousands of GPU threads at once, using CUDA's C++-based extensions. Frameworks like PyTorch compile the tensor operations in a neural network down into optimized CUDA kernels, so a researcher writing normal-looking Python code ends up running highly parallel GPU code without writing CUDA directly.
A well-written, fused CUDA kernel (like Flash Attention) can be several times faster than a naive implementation of the same math, real, measurable inference speed that has nothing to do with model architecture at all.