// the find
prometheus/alertmanager
Prometheus Alertmanager
Alertmanager is the routing and deduplication layer that sits downstream of Prometheus (or anything else that posts to its API) and decides who actually gets paged, emailed, or Slacked when alerts fire. It's for teams already running Prometheus-based monitoring who've outgrown simple one-receiver setups and need grouping, routing, and inhibition logic to keep alert volume sane.
The routing tree (matchers, label-based grouping, inherited config with overridable sub-routes) can model genuinely complex team/escalation structures without bolting on external logic. Clustering for HA is gossip-based (memberlist) and built in rather than an afterthought, and it's been running in production at scale for years. Inhibition rules — muting a warning when the same alertname is already firing critical — solve a real alert-fatigue problem most systems leave to the user. amtool ships alongside it and lets you query, silence, and test routes against a live instance instead of guessing from the YAML.
Config is one flat YAML tree with no include/compose mechanism, so multi-team setups become a single sprawling file where mistakes are easy — the README itself calls out a footgun where an inhibition rule silently applies if the `equal` labels are missing from both sides. Clustering needs both UDP and TCP open between peers, which is an easy miss in containerized or firewalled environments, and a misconfigured cluster degrades quietly rather than erroring loudly. Receiver integrations (Slack, PagerDuty, Opsgenie, etc.) are hardcoded in Go with no plugin system, so anything not on that list has to go through the generic webhook receiver and reimplement its own formatting. There's no way to dry-run a routing change against real historical alerts before deploying it — you're testing against amtool's synthetic queries or trusting the config by hand.