// the find
KeithGalli/complete-pandas-tutorial
A comprehensive tutorial on the Python Pandas library, updated to be consistent with best practices and features available in 2024.
A single Jupyter notebook that's the companion material to Keith Galli's pandas YouTube tutorial, walking through core operations (loading, filtering, merging, groupby, pivoting) using Olympics and coffee-sales datasets. It's for someone learning pandas by following along with a video, not a reference library or tool.
Uses multi-file, realistic datasets (bios/results/noc_regions that need joining) instead of toy single-column frames, so the merge and filter examples actually resemble real work. Includes a markdown cheat sheet as a standalone quick-reference separate from the notebook. Covers the PyArrow backend for read_csv, which most beginner tutorials skip entirely.
It's one big notebook with no exercises or way to self-check beyond watching the video — there's no separation between 'tutorial' and 'reference' content. Fork count (1594) massively outstrips stars (366), which just means a lot of people cloned it to follow along once; don't read this as an active project with ongoing maintenance. The code leans heavily on inplace=True (fillna, dropna, drop, rename), a pattern pandas has been moving away from and that already throws warnings on newer pandas versions — teaches a habit that will need unlearning. No CI, no requirements pinning beyond a flat requirements.txt, and it's already over a year stale as of this evaluation.