finds.dev← search

// the find

stitchfix/hamilton

★ 861 · Python · BSD-3-Clause-Clear · updated Jul 2023

A scalable general purpose micro-framework for defining dataflows. THIS REPOSITORY HAS BEEN MOVED TO www.github.com/dagworks-inc/hamilton

Hamilton is a Python micro-framework where you write plain functions and it assembles them into a DAG by matching parameter names to other function names — no explicit pipeline wiring. It's aimed at data scientists and ML engineers who want testable, pandas/dask/ray/spark-agnostic feature and ETL pipelines. This specific repo is dead: it's a redirect stub, the real project lives at DAGWorks-Inc/hamilton.

The core idea is genuinely clever — pure functions become graph nodes and dependencies are inferred from argument names, so you get a DAG without writing a single edge by hand. It's not tied to one dataframe library; graph adapters let the same function definitions run on pandas, dask, ray, or spark. Pandera integration gives you data quality checks as decorators on the same functions, not a bolted-on separate system. Docs and examples directory are unusually thorough for a project this size.

This repo itself is abandoned — last push mid-2023, README is a one-line 'we moved' notice, nothing here reflects current state. Name-based dependency resolution is implicit magic: rename a function or a parameter and you get a runtime graph error, not a compile-time one, which gets painful past a few dozen nodes. Ray, dask, and spark support live in an 'experimental' folder, so adopt those at your own risk. 861 stars after years of public existence suggests limited traction outside Stitch Fix itself, even accounting for the later rebrand.

View on GitHub → Homepage ↗

// want more like this?

We dig through GitHub every week and send a few repos picked for what you actually care about — each with an honest take like this one.

Get finds in your inbox → Search again →