finds.dev← search

// the find

jeffheaton/jh-kaggle-util

★ 280 · Python · Apache-2.0 · updated Aug 2020

Jeff Heaton's Kaggle Utilities

A personal grab-bag of helper scripts Jeff Heaton built up across three Kaggle competitions, covering data loading/joining, training wrappers for XGBoost/LightGBM/sklearn/Keras, permutation importance, and GLM-based ensembling. Useful mainly as a template for someone setting up their own competition pipeline in a similar style, not as a general-purpose library.

Covers the full competition loop end to end — loading, joining, training across four different model backends, permutation importance, and ensembling — rather than just one slice of it. The two included examples (Mercedes, Santander) show the actual config/data/join/models/ensemble file layout used in real competitions, which is more instructive than API docs would be for this kind of tool.

Dead since 2020 — five-plus years with no commits, no issues addressed, and no adaptation to anything in the modern stack (current XGBoost/LightGBM/sklearn APIs have moved on). The README itself admits it's unfinished ("more instructions coming soon") and the code is explicitly regression-only, so it's a non-starter for classification problems without rewriting the eval/scoring logic. No setup.py or packaging metadata in the tree despite the nested jhkaggle/jhkaggle layout, so it's not pip-installable as-is. No tests, so you're trusting undocumented code that was tuned to one person's specific competition workflow.

View on GitHub →

// want more like this?

We dig through GitHub every week and send a few repos picked for what you actually care about — each with an honest take like this one.

Get finds in your inbox → Search again →