// the find
sustainable-computing-io/kepler
Kepler (Kubernetes-based Efficient Power Level Exporter) is a Prometheus exporter that measures energy consumption metrics at the container, pod, and node levels in Kubernetes clusters.
Kepler is a Prometheus exporter that estimates power draw per container, pod, and node in a Kubernetes cluster by reading RAPL/powercap counters (with experimental HWMon, NVIDIA GPU, and Redfish BMC support). It's aimed at platform teams trying to put real numbers behind cluster energy/cost reporting, not app developers.
The 0.10 rewrite dropped CAP_SYSADMIN/CAP_BPF and now runs with just readonly /proc and /sys access, which matters a lot for something that has to run as a privileged DaemonSet on every node. Power attribution is now based on active CPU usage per workload instead of a crude idle/dynamic split, which was the biggest accuracy complaint about the old version. The project also holds itself to a real bar — OpenSSF Scorecard and Best Practices badges, codecov, CI across the usual matrix — rather than just shipping and hoping.
RAPL/powercap is a bare-metal/physical-CPU feature — on most managed cloud Kubernetes nodes (EKS, GKE, AKS) it simply isn't exposed, so the exporter's primary data source doesn't work where a lot of clusters actually run; this isn't called out until you're deep in the README. The rewrite is a hard break: 0.9.x is now frozen with no bug fixes, so anyone not ready to migrate is stuck on an unmaintained line. GPU power monitoring is NVIDIA-only, which is a real gap given how much cluster power now goes to GPUs. The README itself warns that `main` manifests can reference endpoints missing from the last tagged image and crash-loop the pod if you mix them — a rough edge that shouldn't need a warning in the first place.