// the find
rhiever/sklearn-benchmarks
A centralized repository to report scikit-learn model performance across a variety of parameter settings and data sets.
A 2017 research artifact from the Moore/Olson bioinformatics lab that ran grid and random search over a dozen scikit-learn classifiers against the PMLB benchmark datasets, producing the data behind their 'data-driven advice for ML in bioinformatics' paper. It's for someone who wants the raw numbers behind that paper, not a tool you install and use.
The per-algorithm script layout (one file per classifier, split across grid_search/random_search/random_search_preprocessing) makes it easy to see exactly what hyperparameter ranges were swept for each model. It's tied to a real published paper with a citable arXiv reference, so the methodology isn't just a black box. The metafeatures module computing dataset characteristics (and pairing them with PMLB) is a reasonable building block for meta-learning experiments if you gut it out of the rest of the repo.
Dead since October 2017 — no requirements.txt or environment pin, and scikit-learn's API (especially around GridSearchCV, model defaults, and several of these classifiers) has moved on enough that these scripts likely won't run as-is. There's no setup.py, no CLI, no way to actually reproduce a result without reading through notebooks and reconstructing the pipeline by hand. Dataset access is outsourced to a separate PMLB repo with no version pinned, so the exact data used in 2017 may not match what you'd download today. Zero tests beyond one file for a metafeatures helper, no CI, and the notebooks (analyze-sklearn-benchmark6.ipynb etc.) suggest iterative exploration rather than a maintained, reproducible pipeline.