// the find
jeffheaton/jh-kaggle-util
Jeff Heaton's Kaggle Utilities
A personal grab-bag of helper scripts Jeff Heaton built up across three Kaggle competitions, covering data loading/joining, training wrappers for XGBoost/LightGBM/sklearn/Keras, permutation importance, and GLM-based ensembling. Useful mainly as a template for someone setting up their own competition pipeline in a similar style, not as a general-purpose library.
Covers the full competition loop end to end — loading, joining, training across four different model backends, permutation importance, and ensembling — rather than just one slice of it. The two included examples (Mercedes, Santander) show the actual config/data/join/models/ensemble file layout used in real competitions, which is more instructive than API docs would be for this kind of tool.
Dead since 2020 — five-plus years with no commits, no issues addressed, and no adaptation to anything in the modern stack (current XGBoost/LightGBM/sklearn APIs have moved on). The README itself admits it's unfinished ("more instructions coming soon") and the code is explicitly regression-only, so it's a non-starter for classification problems without rewriting the eval/scoring logic. No setup.py or packaging metadata in the tree despite the nested jhkaggle/jhkaggle layout, so it's not pip-installable as-is. No tests, so you're trusting undocumented code that was tuned to one person's specific competition workflow.