finds.dev← search

// the find

rhiever/sklearn-benchmarks

★ 213 · Jupyter Notebook · MIT · updated Oct 2017

A centralized repository to report scikit-learn model performance across a variety of parameter settings and data sets.

A 2017 research artifact from the Moore/Olson bioinformatics lab that ran grid and random search over a dozen scikit-learn classifiers against the PMLB benchmark datasets, producing the data behind their 'data-driven advice for ML in bioinformatics' paper. It's for someone who wants the raw numbers behind that paper, not a tool you install and use.

The per-algorithm script layout (one file per classifier, split across grid_search/random_search/random_search_preprocessing) makes it easy to see exactly what hyperparameter ranges were swept for each model. It's tied to a real published paper with a citable arXiv reference, so the methodology isn't just a black box. The metafeatures module computing dataset characteristics (and pairing them with PMLB) is a reasonable building block for meta-learning experiments if you gut it out of the rest of the repo.

Dead since October 2017 — no requirements.txt or environment pin, and scikit-learn's API (especially around GridSearchCV, model defaults, and several of these classifiers) has moved on enough that these scripts likely won't run as-is. There's no setup.py, no CLI, no way to actually reproduce a result without reading through notebooks and reconstructing the pipeline by hand. Dataset access is outsourced to a separate PMLB repo with no version pinned, so the exact data used in 2017 may not match what you'd download today. Zero tests beyond one file for a metafeatures helper, no CI, and the notebooks (analyze-sklearn-benchmark6.ipynb etc.) suggest iterative exploration rather than a maintained, reproducible pipeline.

View on GitHub →

// want more like this?

We dig through GitHub every week and send a few repos picked for what you actually care about — each with an honest take like this one.

Get finds in your inbox → Search again →