finds.dev← search

// the find

karpathy/nn-zero-to-hero

★ 24,513 · Jupyter Notebook · MIT · updated Aug 2024

Neural Networks: Zero to Hero

This is the notebook repo backing Karpathy's YouTube series that builds neural nets from scratch, going from a scalar autograd engine (micrograd) up through a bigram model, MLPs, batchnorm, manual backprop, and finally a GPT and its BPE tokenizer. It's for people who want to understand what's actually happening inside a transformer rather than just calling `.fit()` on one.

Nothing is imported as a black box — the autograd engine, the tokenizer, and attention are all written by hand so you see every gradient and every tensor reshape. The lecture sequence is genuinely well ordered pedagogically: each notebook's pain point (exploding gradients, dead tanh units, manual backprop) motivates the next lecture's fix. The backprop-ninja notebook in particular is rare — most courses never make you manually derive gradients through batchnorm and cross-entropy by hand.

The notebooks are a companion artifact to the videos, not a standalone reference — without watching the corresponding lecture, several notebooks are sparse on inline explanation and hard to follow cold. The repo is effectively frozen (last push mid-2024, marked 'Ongoing...' with no further lectures shipped), so anything after GPT-2/tokenization — RLHF, MoE, longer-context tricks — isn't covered. There's no requirements.txt or environment pinning, so PyTorch/version drift can silently break older notebooks. And since it's a teaching artifact rather than a library, there's no packaging, tests, or CI — you're expected to read and retype code, not import it.

View on GitHub →

// want more like this?

We dig through GitHub every week and send a few repos picked for what you actually care about — each with an honest take like this one.

Get finds in your inbox → Search again →