finds.dev← search

// the find

mrdbourke/simple-local-rag

★ 1,021 · Jupyter Notebook · updated May 2024

Build a RAG (Retrieval Augmented Generation) pipeline from scratch and have it all run locally.

A single Jupyter notebook walking through building a RAG pipeline from scratch — chunking a PDF, embedding it with sentence-transformers, storing vectors as a torch.tensor, and generating answers with a local Gemma model on an NVIDIA GPU. It's a tutorial for people who want to understand what LangChain and friends are doing under the hood, not a library to import into a project.

Builds every step by hand (chunking, embedding, similarity search, prompt construction) instead of hiding it behind a framework, so you actually see where the RAG pipeline's moving parts are. Uses a concrete worked example (a real 1200-page PDF) rather than a three-sentence toy doc. Storing embeddings as a plain torch.tensor and doing similarity search with matrix multiplication is a legitimate choice at this scale — no need to drag in a vector DB for under 100k chunks. There's a full video walkthrough synced to the code for people who learn better watching someone type it out.

It's one notebook, not a package — there's nothing to pip install and import, you're expected to copy and adapt the cells yourself. Hard dependency on an NVIDIA GPU with 5GB+ VRAM (Colab is the fallback, which has its own quota headaches). The author's own README TODO list is still unchecked after over a year — setup instructions are flagged incomplete and the 'Extensions' section is a stub that says 'coming soon' and never arrived. No commits since May 2024, so it predates two years of changes to the Gemma, transformers, and torch APIs it depends on; don't be surprised if requirements.txt no longer resolves cleanly.

View on GitHub →

// want more like this?

We dig through GitHub every week and send a few repos picked for what you actually care about — each with an honest take like this one.

Get finds in your inbox → Search again →