// the find
karpathy/nanochat
The best ChatGPT that $100 can buy.
nanochat is Karpathy's minimal, single-file-per-stage training harness for building a ChatGPT-style model from scratch on one 8xGPU node, covering tokenization through RLHF-style finetuning and inference. It's aimed at researchers and hobbyists who want to understand or hack on the full LLM pipeline rather than fine-tune an existing checkpoint, not at anyone trying to build a production model.
The single `--depth` knob deriving width, heads, LR schedule, and training horizon automatically is a genuinely good design choice — it removes the usual pile of interacting hyperparameters and forces changes to be principled across model scales. Explicit `COMPUTE_DTYPE` handling instead of `torch.amp.autocast` gives real control over precision per-layer without the usual autocast surprises. The speedrun script plus public leaderboard (wall-clock time to beat GPT-2's CORE score) is a great forcing function for reproducibility and turns training into something you can actually benchmark against others.
It's hard-tied to a single 8xH100/A100 node topology; multi-node training isn't part of the design, so this doesn't scale past one box no matter your budget. float16 training gets a GradScaler in base_train and SFT but RL doesn't, so mixed-precision support is inconsistent across the pipeline. CPU/MPS is explicitly a toy path that 'will not get strong results' — fine for poking at the code, useless for anything real. The author admits several hardware paths (xpu, some MPS configs) are untested and may have sharp edges, and the codebase intentionally avoids configurability, so adapting it to a different data pipeline or architecture means editing core files rather than passing options.