// the find
karpathy/minGPT
A minimal PyTorch re-implementation of the OpenAI GPT (Generative Pretrained Transformer) training
minGPT is Karpathy's minimal PyTorch reimplementation of GPT-2/GPT-3 style transformers, kept to a few hundred lines across model, BPE tokenizer, and trainer. It's for people who want to read and understand exactly how a GPT works end to end, not for anyone who wants to train a real model at scale.
The model definition is genuinely ~300 lines and readable in one sitting, which is rare for transformer code — no config-driven branching or framework abstraction hiding the attention math. The BPE tokenizer is a faithful from-scratch reimplementation of GPT-2's encoding rather than a wrapper around a library, so you can actually see how text becomes token ids. The projects/adder and projects/chargpt examples give concrete, runnable end-to-end tasks instead of just a bare model class.
The repo has been semi-archived since Jan 2023 and the author explicitly redirects to nanoGPT for anything beyond education — there's no distributed training, no mixed precision, and the TODO list in the README is basically unaddressed. There's no requirements.txt (the author admits this in the README), so pinning dependencies is on you. Test coverage is minimal (one unittest module), and it can only load pretrained weights for the gpt2-* family, not arbitrary checkpoints.