finds.dev← search

// the find

karpathy/llama2.c

★ 20,106 · C · MIT · updated Aug 2024

Inference Llama 2 in one file of pure C

A single-file C implementation for running inference on Llama 2 architecture models, paired with a minimal PyTorch training script. Built by Karpathy as a teaching artifact for people who want to understand how a transformer forward pass actually works without wading through a production inference engine.

run.c is genuinely ~700 lines of readable C with zero dependencies — you can read the whole forward pass in an afternoon and know exactly what's happening at every matmul. The int8 quantized path (runq.c) is a good worked example of activation quantization, not just weight quantization, and the README explains the tradeoff honestly. Ships pretrained TinyStories checkpoints (260K to 110M params) so you get a working demo in under a minute without training anything yourself.

Inference is fp32-only for anything beyond the toy TinyStories models, so loading a real 7B checkpoint means a 26GB file and single-digit tokens/sec on CPU — this is explicitly not meant for production use. 13B+ models are broken due to integer overflow in pointer arithmetic and the maintainer says so plainly rather than pretending it works. Last meaningful commit was mid-2024, so it predates Llama 3/3.1 architecture changes (GQA is only partially represented via n_kv_heads) and there's no KV-cache-aware batching or GPU path beyond a stalled CUDA port mentioned in the todos.

View on GitHub →

// want more like this?

We dig through GitHub every week and send a few repos picked for what you actually care about — each with an honest take like this one.

Get finds in your inbox → Search again →