CUDA

AI Infrastructure Infrastructure

CUDA is NVIDIA's programming platform for writing software that runs directly on its GPUs, and it's the foundation nearly every AI training and inference framework is built on top of.

In simple terms

A GPU can't run arbitrary code the way a CPU can; it needs code written specifically for its parallel architecture. CUDA is the toolkit and language extensions that let developers write that code, and it's what libraries like PyTorch use underneath to actually make an NVIDIA GPU do the math a neural network needs.

Why it matters

CUDA's decade-plus head start and deep integration into every major AI framework is a large part of why NVIDIA GPUs dominate AI infrastructure; competing hardware has to either support CUDA or convince the entire software ecosystem to retarget itself.

How it works

Developers write kernels, small functions meant to run on thousands of GPU threads at once, using CUDA's C++-based extensions. Frameworks like PyTorch compile the tensor operations in a neural network down into optimized CUDA kernels, so a researcher writing normal-looking Python code ends up running highly parallel GPU code without writing CUDA directly.

Where it fits

PyTorch / framework code → Compiled CUDA kernels → GPU cores → Result tensors

Production impact

A well-written, fused CUDA kernel (like Flash Attention) can be several times faster than a naive implementation of the same math, real, measurable inference speed that has nothing to do with model architecture at all.

Learn this first

Related terms

Learn next

← All terms · Knowledge map →