// the find
m3dev/gokart
Gokart solves reproducibility, task dependencies, constraints of good code, and ease of use for Machine Learning Pipeline.
gokart is m3's opinionated wrapper around Luigi that adds reproducibility guarantees to ML pipelines — every task's output, parameters, module versions, and even random seed get hashed and cached to a pkl file, so reruns are automatic when inputs change and skipped when they don't. It's aimed at teams running batch ML pipelines in Python who want Makefile-style caching without hand-rolling it, and who are fine writing config as Python classes rather than YAML.
The parameter-hash-based caching is legitimately useful: change a hyperparameter and only the affected downstream tasks rerun, with S3/GCS support baked in for storing intermediates. The mypy plugin plus `TaskOnKart[T]`/`TaskInstanceParameter` generics give you real type-checked task graphs, which is rare in this space and catches wiring mistakes (task A dumps int, task B expects str) at typecheck time instead of at runtime three hours into a batch job. Redis-based locking for conflict prevention across parallel task instances is a practical detail most in-house pipeline wrappers skip. Three years of production use at m3 plus a documented external competition win is more real-world validation than most pipeline-framework repos this size have.
It's built on Luigi, which means you inherit Luigi's single-process central scheduler model — there's no distributed execution story here, and the README itself admits parallel execution is process-level, not concurrent-in-memory. The project explicitly punts on experiment tracking (you need their separate Thunderbolt repo) and has zero task visualization, so debugging a large DAG means reading logs, not looking at a graph. Every intermediate result gets serialized to disk/S3/GCS as pkl by design, so for pipelines with large intermediate tensors or dataframes, I/O becomes the bottleneck exactly as the README warns — there's no in-memory-only mode for tight iteration loops. Docs are readthedocs/rst rather than runnable end-to-end examples beyond the notebook, so adopting it for anything non-trivial means reading source (task.py, worker.py) to understand lock and cache-key semantics.