finds.dev← search

// the find

tensorzero/tensorzero

★ 11,715 · Rust · Apache-2.0 · updated Jun 2026

TensorZero is an open-source LLMOps platform that unifies an LLM gateway, observability, evaluation, optimization, and experimentation.

TensorZero is a self-hosted LLMOps stack written in Rust that bundles an LLM gateway, observability storage, evaluation tooling, and prompt/model optimization into one system, fronted by an OpenAI-compatible API. It's aimed at teams running LLM features in production who want routing, fallbacks, and a feedback loop from real usage back into fine-tuning, rather than hand-rolling that plumbing across several vendors.

The gateway is a drop-in OpenAI-compatible endpoint, so you can point an existing client at it without rewriting application code, and the crate layout (durable-tools, autopilot-tools, evaluations, config-applier) shows actual engineering behind the batch inference, retry, and routing logic rather than a thin proxy. The Rust implementation backs up its latency claims with a real benchmark suite under crates/gateway/benchmarks. Config is TOML and GitOps-friendly, and there's a documented escape hatch to the underlying database for anything the API doesn't expose. CI is substantial — dozens of workflows covering live-tests, ClickHouse, Helm publishing, and CodeQL — which signals a real release process, not a side project.

The 'incrementally adopt' pitch undersells how much infrastructure the full story requires: ClickHouse plus Postgres plus the gateway plus the UI plus optimization workers, which is a lot to run just to get observability and a feedback loop. The headline feature in the README, Autopilot ('automated AI engineer'), is actually a separate paid product layered on top, not something you get from cloning this repo — that's a bait-and-switch for anyone reading the top of the page. The README spends more space on funding announcements and a Fortune-10 claim than on what a minimal production deployment actually costs to operate (ClickHouse maintenance, config complexity across variants/functions). And despite the 'take what you need' framing, the data model (functions, variants, episodes) means understanding any one piece, like GEPA optimization or DICL, requires first understanding the whole schema.

View on GitHub → Homepage ↗

// want more like this?

We dig through GitHub every week and send a few repos picked for what you actually care about — each with an honest take like this one.

Get finds in your inbox → Search again →