// the find
karpathy/llm.c
LLM training in simple, raw C/CUDA
llm.c is Karpathy's from-scratch GPT-2/GPT-3 training implementation in plain C and CUDA, built to strip away the PyTorch/Python stack entirely. It's aimed at people who want to understand exactly what happens during LLM training at the kernel level, not at people who want to train a production model quickly.
The dev/cuda folder is a genuinely useful teaching resource — hand-written kernels for every layer (layernorm, attention, matmul, adamw) ranging from naive to near-cuBLAS speed, so you can see the actual performance gap instead of just reading about it. The CPU fp32 reference in train_gpt2.c is ~1000 lines and mirrors train_gpt2.py closely enough that you can diff behavior between the two. There's a real correctness harness (test_gpt2.c/cu) that checks C output against PyTorch logits and gradients, not just 'it runs'. Multi-GPU and multi-node support (NCCL, MPI, three different init strategies) is more complete than most educational projects bother with.
This is explicitly a root-folder-simplicity-over-features project, so anything beyond GPT-2/GPT-3 architecture (MoE, different attention variants, newer positional encodings) isn't in scope and won't be added per the author's own stated philosophy. cuDNN flash attention is opt-in and adds a ~1 minute compile time, which is a rough first-run experience if you don't already have cuDNN installed. Last push was June 2025, and the project reads as feature-complete/frozen rather than actively maintained — most ongoing work has migrated to the many third-party ports listed in the README rather than this repo itself. No packaging or pip-installable path exists by design; you're building Makefiles and managing CUDA/NCCL versions yourself, which is friction if you just want to reproduce a run rather than study the internals.