// the find
Netflix/maestro
Maestro: Netflix’s Workflow Orchestrator
Maestro is Netflix's internal workflow orchestrator, now open-sourced, for scheduling and running data/ML/ETL pipelines. It's aimed at teams operating at real scale (Netflix runs hundreds of thousands of workflows and millions of jobs a day on it), not solo developers looking for a lightweight cron replacement.
Production-proven at a scale most orchestrators never see, with public writeups on the specific engineering problems they hit (a 100x workflow-engine speedup, Iceberg-based incremental processing) rather than vague scale claims. It ships pluggable execution backends out of the box — AWS (SQS/SNS via LocalStack for local dev), Kubernetes, Redis for step concurrency — plus a separate maestro-extensions service for things like foreach-step flattening, showing a real extension point rather than a monolith you have to fork. There's also a Python SDK for defining and pushing workflows without touching the Java API directly.
This is a Java/Spring Boot/Gradle system with AWS and Kubernetes as first-class dependencies — getting a local instance running means Docker, LocalStack, and kubectl config just to try the sample DAG, which is a lot of ceremony for evaluation. It's Netflix-shaped: SQS/SNS are baked into the AWS module rather than treated as one of several pluggable queue backends, so teams on GCP or Azure, or without a message-bus-centric setup already, will be swimming upstream. The in-repo README leans on links to Netflix Tech Blog posts for anything architectural, so understanding how the scheduler, foreach handling, or the extensions service actually work means leaving the repo rather than reading docs alongside the code.