finds.dev← search

// the find

KeithGalli/Pandas-Data-Science-Tasks

★ 1,019 · Jupyter Notebook · updated Nov 2023

Set of real world data science tasks completed using the Python Pandas library

A single Jupyter notebook that walks through cleaning and analyzing a year of synthetic electronics-store sales data with pandas and matplotlib, built as the companion material for a KeithGalli YouTube tutorial. It's aimed at people learning pandas basics — concat, groupby, string parsing, apply — not at anyone looking for a reusable tool or library.

The dataset is realistic enough to force actual cleaning work (NaNs, mixed types, string columns that need splitting) rather than a toy iris/titanic set. Each pandas operation is tied to a concrete business question (best month, best city, products bought together) instead of being demonstrated in isolation, which makes the 'why' obvious. It's fully self-contained — monthly CSVs are checked into the repo along with a script to regenerate them, so there's zero setup friction or external API dependency to reproduce the results.

There's exactly one notebook and no written explanation in the repo itself — the actual teaching happens in the linked video, so the repo alone is a skeleton without much value if you don't watch it. No requirements.txt or environment file, so pinning down which pandas/matplotlib versions this was built against is left to the reader. The 2924 forks vs 1019 stars tells you what this really is: a clone-and-follow-along exercise, not a project anyone extends — and it hasn't been touched since late 2023, so any pandas API drift (chained assignment warnings, groupby defaults) won't be fixed.

View on GitHub →

// want more like this?

We dig through GitHub every week and send a few repos picked for what you actually care about — each with an honest take like this one.

Get finds in your inbox → Search again →