finds.dev← search

// the find

karpathy/nanoGPT

★ 63,343 · Python · MIT · updated Nov 2025

The simplest, fastest repository for training/finetuning medium-sized GPTs.

nanoGPT is Karpathy's minimal GPT training/finetuning codebase — a ~300-line model.py and ~300-line train.py that reproduce GPT-2 (124M) from scratch on a single 8xA100 node. It's for people who want to actually read and modify every line of a GPT training loop rather than call into a framework, whether that's finetuning on Shakespeare on a laptop or reproducing GPT-2 baselines on a cluster.

The code is genuinely as small as advertised — model.py and train.py are each readable in one sitting, which makes it a real teaching tool rather than a toy wrapped around hidden complexity. It supports the full range from CPU/MPS toy runs to multi-node DDP training, with configs that make the jump between them explicit rather than magic. Loading pretrained GPT-2 checkpoints via transformers and finetuning them with the same script as from-scratch training is a nice bit of design economy.

As of Nov 2025 the author has flagged this repo as deprecated in favor of nanochat and says he's leaving it up only for posterity — there's no reason to build new work on it going forward. torch.compile is on by default and doesn't work on Windows without manually passing --compile=False, which is exactly the kind of thing that eats someone's first hour. The todos list (FSDP, zero-shot eval, better logging) has sat unaddressed for years, so don't expect production training features like fault-tolerant checkpointing or modern parallelism strategies.

View on GitHub →

// want more like this?

We dig through GitHub every week and send a few repos picked for what you actually care about — each with an honest take like this one.

Get finds in your inbox → Search again →