// the find
tensorzero/tensorzero
TensorZero is an open-source LLMOps platform that unifies an LLM gateway, observability, evaluation, optimization, and experimentation.
TensorZero is a self-hosted LLMOps stack written in Rust that bundles an LLM gateway, observability storage, evaluation tooling, and prompt/model optimization into one system, fronted by an OpenAI-compatible API. It's aimed at teams running LLM features in production who want routing, fallbacks, and a feedback loop from real usage back into fine-tuning, rather than hand-rolling that plumbing across several vendors.
The gateway is a drop-in OpenAI-compatible endpoint, so you can point an existing client at it without rewriting application code, and the crate layout (durable-tools, autopilot-tools, evaluations, config-applier) shows actual engineering behind the batch inference, retry, and routing logic rather than a thin proxy. The Rust implementation backs up its latency claims with a real benchmark suite under crates/gateway/benchmarks. Config is TOML and GitOps-friendly, and there's a documented escape hatch to the underlying database for anything the API doesn't expose. CI is substantial — dozens of workflows covering live-tests, ClickHouse, Helm publishing, and CodeQL — which signals a real release process, not a side project.
The 'incrementally adopt' pitch undersells how much infrastructure the full story requires: ClickHouse plus Postgres plus the gateway plus the UI plus optimization workers, which is a lot to run just to get observability and a feedback loop. The headline feature in the README, Autopilot ('automated AI engineer'), is actually a separate paid product layered on top, not something you get from cloning this repo — that's a bait-and-switch for anyone reading the top of the page. The README spends more space on funding announcements and a Fortune-10 claim than on what a minimal production deployment actually costs to operate (ClickHouse maintenance, config complexity across variants/functions). And despite the 'take what you need' framing, the data model (functions, variants, episodes) means understanding any one piece, like GEPA optimization or DICL, requires first understanding the whole schema.