// the find
opsre/WatchAlert
🚀A lightweight cloud-native multi-data source monitoring and alerting engine |一款轻量级云原生多数据源监控告警引擎
WatchAlert is a Go-based monitoring and alerting engine aimed at cloud-native shops that want one system covering metrics (Prometheus), logs (Loki, ES, ClickHouse, VictoriaLogs, cloud vendor log services), traces (Jaeger), Kubernetes events, and uptime probing, instead of stitching together Alertmanager plus separate tools for each signal type. It's built by a Chinese team/community (opsre, formerly w8t-io) and ships a React/Ant Design frontend alongside the Go backend.
Genuinely broad datasource coverage under one alerting pipeline — metrics, logs, traces, k8s events and network probes (HTTP/ICMP/TCP/SSL) all route through the same rule/notify path, which is more than most self-hosted alternatives attempt. On-call duty scheduling and multi-level alert escalation (timeout retry, handoff to next responder) are first-class features, not bolted on, which matters since most open-source alerting tools punt that to PagerDuty/Opsgenie. Multi-tenancy is modeled properly (dedicated tenant middleware, tenant-linked users, per-tenant SQL seed files) rather than being a single-org afterthought. There's a dedicated SSRF guard package with its own test (pkg/tools/ssrf.go) — notable given this tool accepts user-configured webhook/datasource URLs, a common SSRF vector that most similar projects don't bother defending.
Documentation and the primary community channel are Chinese-only; English-speaking adopters get a much thinner picture of setup and gotchas than native users. A SQLite file (data/w8t.db) is checked into the repo tree, which usually means a dev/test database artifact got committed rather than gitignored — worth checking before treating it as a real default. The 'AI-powered' root-cause analysis is described only in marketing terms in the README (no mention of which model, how it's invoked, or what it costs to run) — it reads like a wrapper around a hosted LLM API rather than a described, testable capability. Test coverage is thin and concentrated in a few pkg/ packages (ai_client, oidc, clickhouse, kubernetes, elasticsearch); the core alerting logic in alert/eval and the sizeable api/ and internal/services layers have no visible tests, so regressions in rule evaluation or escalation logic wouldn't be caught by CI.