// the find
spotify/luigi
Luigi is a Python module that helps you build complex pipelines of batch jobs. It handles dependency resolution, workflow management, visualization etc. It also comes with Hadoop support built in.
Luigi is a Python framework for chaining and scheduling long-running batch jobs (Hadoop, Spark, Hive, SQL dumps, etc.) where dependencies are declared as ordinary Python objects instead of XML or a separate DSL. It's built for data engineers running thousands of interdependent batch tasks, especially ones with a foot in Hadoop-era infrastructure.
The Target abstraction enforces atomic file writes (write to a temp path, rename on success), so a crashed pipeline never leaves a downstream task reading half-written data. The contrib package is enormous and actually useful, with working integrations for S3, BigQuery, Redshift, Postgres, Kubernetes, ECS, and Spark that you'd otherwise have to write yourself. Dependency graphs are plain Python — task parameters, requires(), and output() are testable objects, not a YAML or XML layer bolted on top. The central scheduler's visualizer draws a real dependency graph with task state, which is genuinely useful when debugging why job #4000 in a long chain is stuck.
There's no built-in triggering — Luigi resolves and tracks dependencies but won't kick anything off on a schedule, so you still need cron or systemd timers wrapping the luigi command, unlike Airflow where scheduling is native. The central scheduler is a single coordination point with no built-in HA story, which starts to matter once thousands of tasks a day are hitting it. The web UI is a Bootstrap 3 / jQuery admin theme that hasn't been meaningfully redesigned in years — it works, but looks and feels a decade older than Airflow's or Dagster's UI. Newer orchestration ideas like dynamic task generation and asset-based scheduling live elsewhere; Luigi at this point is mostly maintained, not actively extended.