finds.dev← search

// the find

KeithGalli/complete-pandas-tutorial

★ 366 · Jupyter Notebook · MIT · updated Jul 2024

A comprehensive tutorial on the Python Pandas library, updated to be consistent with best practices and features available in 2024.

A single Jupyter notebook that's the companion material to Keith Galli's pandas YouTube tutorial, walking through core operations (loading, filtering, merging, groupby, pivoting) using Olympics and coffee-sales datasets. It's for someone learning pandas by following along with a video, not a reference library or tool.

Uses multi-file, realistic datasets (bios/results/noc_regions that need joining) instead of toy single-column frames, so the merge and filter examples actually resemble real work. Includes a markdown cheat sheet as a standalone quick-reference separate from the notebook. Covers the PyArrow backend for read_csv, which most beginner tutorials skip entirely.

It's one big notebook with no exercises or way to self-check beyond watching the video — there's no separation between 'tutorial' and 'reference' content. Fork count (1594) massively outstrips stars (366), which just means a lot of people cloned it to follow along once; don't read this as an active project with ongoing maintenance. The code leans heavily on inplace=True (fillna, dropna, drop, rename), a pattern pandas has been moving away from and that already throws warnings on newer pandas versions — teaches a habit that will need unlearning. No CI, no requirements pinning beyond a flat requirements.txt, and it's already over a year stale as of this evaluation.

View on GitHub →

// want more like this?

We dig through GitHub every week and send a few repos picked for what you actually care about — each with an honest take like this one.

Get finds in your inbox → Search again →