finds.dev← search

// the find

jakevdp/sklearn_scipy2013

★ 323 · Python · updated Jun 2017

Scikit-learn tutorials for the Scipy 2013 conference

A set of IPython notebooks from Jake VanderPlas, Olivier Grisel, and Gael Varoquaux's scikit-learn tutorial at SciPy 2013, covering classification, regression, clustering, dimensionality reduction, and scaling text classification with SGD. It's aimed at people learning the fundamentals of scikit-learn's API and general ML workflow (train/test split, cross-validation, bias/variance).

The notebooks are written by people who actually built scikit-learn's core APIs, so the explanations of the estimator/fit/predict pattern and the bias-variance framing are accurate and well-sequenced. The exercises come with worked solutions in notebooks/solutions/, which is rare for conference material and useful for self-study. The scope is broad but coherent: it moves from basic representation of data through supervised/unsupervised learning to a real distributed/out-of-core text classification example, which most intro tutorials skip.

Dead project: last commit is from 2017, README still tells you to use Python 2.6/2.7, and it predates most of scikit-learn's current API (Pipeline, ColumnTransformer, and the estimator tags system didn't exist yet). Running it today means fighting environment setup — old numpy/scipy/matplotlib pins, ipython notebook instead of Jupyter — before you even get to the ML content. No tests, no CI, and the data-loading code depends on sklearn's old bundled dataset fetchers, some of which have since moved or changed signatures. As a teaching resource this has been superseded by scikit-learn's own current tutorials and docs, which stay in sync with the library.

View on GitHub →

// want more like this?

We dig through GitHub every week and send a few repos picked for what you actually care about — each with an honest take like this one.

Get finds in your inbox → Search again →