finds.dev← search

// the find

smg-project/smg

★ 543 · Rust · Apache-2.0 · updated Sep 2026

Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing, chat history, tokenization caching, Responses API, embeddings, WASM plugins, MCP, and multi-tenant auth.

SMG is a Rust gateway that sits in front of self-hosted inference engines (vLLM, SGLang, TensorRT-LLM, TokenSpeed, MLX) and cloud APIs (OpenAI, Anthropic, Gemini, xAI), exposing one OpenAI/Anthropic-compatible endpoint with KV-cache-aware routing, chat history, and MCP tool execution. It's aimed at teams running multi-GPU self-hosted inference who need real production routing and observability, not solo single-instance use.

The gRPC pipeline talks to engine wire protocols directly rather than just proxying HTTP — the protocol/vllm and protocol/tokenspeed submodules handle sampling, multimodal input, and structured outputs at the protocol level, which is real integration work, not a thin reverse proxy. KV-cache-aware routing via radix trees is backed by a dedicated fuzz workflow (radix-tree-fuzz.yml) and nightly benchmarks (tokenizer, tool-parser, BFCL function-calling eval), so routing quality claims are actually being tested rather than just asserted. The client story is unusually complete for a gateway: generated OpenAPI clients plus hand-maintained Rust/Python/Go bindings each with their own test suites.

The feature surface is enormous — 10 routing policies, WASM plugins, SWIM/CRDT mesh HA, OIDC auth, Oracle-backed chat history, MCP tool execution, three language bindings — which is a lot for a 543-star project to keep working across every combination; less-common paths (Oracle history, WASM plugins, mesh HA) are probably far less battle-tested than the default single-node HTTP case. Building from source needs protoc and pulls in gRPC/protobuf codegen, ZMQ, and multiple DB drivers into one binary, so even a simple round-robin proxy in front of one vLLM instance drags in machinery you won't use. It's tightly coupled to specific engine versions (per-engine CI setup actions for vLLM/SGLang/TRT-LLM, a pinned tokenspeed.ref), meaning upstream engine upgrades can silently break the gRPC contract until SMG catches up.

View on GitHub → Homepage ↗

// want more like this?

We dig through GitHub every week and send a few repos picked for what you actually care about — each with an honest take like this one.

Get finds in your inbox → Search again →