finds.dev← search

// the find

cruise-control-for-kafka/cruise-control

★ 3,044 · Java · Apache-2.0 · updated Oct 2026

Cruise-control is the first of its kind to fully automate the dynamic workload rebalance and self-healing of a Kafka cluster. It provides great value to Kafka users by simplifying the operation of Kafka clusters.

Cruise Control runs alongside an Apache Kafka cluster, samples broker and partition load, and builds a model that generates rebalance proposals and detects anomalies such as broker failures and goal violations. It can fix some of those automatically. It is for teams whose clusters are large enough that moving partitions by hand has become a recurring job.

The goals are an explicit, ordered list. Rack awareness and capacity limits sit above the distribution goals, so a balance pass cannot quietly break rack placement, and custom goals load from a jar passed with -jars instead of requiring a fork. The sample store is pluggable, and the README gets the reasoning right: derived metrics depend on cluster metadata at the time they were collected, so replaying raw metrics later gives wrong numbers. Broker failure handling has a check stage that waits a configurable grace period before moving replicas, so a broker that reboots in two minutes does not trigger a full reassignment. Self-healing is opt-in per detector, which is the right default for software that moves production data.

Adoption is invasive. You build the metrics reporter, copy its jar into every broker's libs folder, set metric.reporters in server.properties, and roll every broker. The reporter also writes to its own metrics topic, which adds load, and the README has to warn you that a compacted cleanup policy on that topic must be changed to delete, so it is one more thing to get wrong. Kafka support is split across branches: main for 2.5+, migrate_to_kafka_2_4, and two deprecated branches for older versions. Disk failure detection, slow broker detection, and the intra-broker goals are missing on the 0.11 and 1.0 branch, so you have to match the release line to your exact Kafka version. The default capacity file is a placeholder that the README says may not match your brokers, yet every proposal is computed against the capacities you give it, and a plan built on placeholder numbers looks just as authoritative as one built on real ones. The README also defers configuration and REST API detail to the wiki, so checking what a config key actually does often means reading the source.

View on GitHub → Homepage ↗

// want more like this?

We dig through GitHub every week and send a few repos picked for what you actually care about — each with an honest take like this one.

Get finds in your inbox → Search again →