// the find
karpathy/nanoGPT
The simplest, fastest repository for training/finetuning medium-sized GPTs.
nanoGPT is Karpathy's minimal GPT training/finetuning codebase — a ~300-line model.py and ~300-line train.py that reproduce GPT-2 (124M) from scratch on a single 8xA100 node. It's for people who want to actually read and modify every line of a GPT training loop rather than call into a framework, whether that's finetuning on Shakespeare on a laptop or reproducing GPT-2 baselines on a cluster.
The code is genuinely as small as advertised — model.py and train.py are each readable in one sitting, which makes it a real teaching tool rather than a toy wrapped around hidden complexity. It supports the full range from CPU/MPS toy runs to multi-node DDP training, with configs that make the jump between them explicit rather than magic. Loading pretrained GPT-2 checkpoints via transformers and finetuning them with the same script as from-scratch training is a nice bit of design economy.
As of Nov 2025 the author has flagged this repo as deprecated in favor of nanochat and says he's leaving it up only for posterity — there's no reason to build new work on it going forward. torch.compile is on by default and doesn't work on Windows without manually passing --compile=False, which is exactly the kind of thing that eats someone's first hour. The todos list (FSDP, zero-shot eval, better logging) has sat unaddressed for years, so don't expect production training features like fault-tolerant checkpointing or modern parallelism strategies.