finds.dev← search

// the find

huggingface/text-embeddings-inference

★ 5,064 · Rust · Apache-2.0 · updated Sep 2026

A blazing fast inference solution for text embeddings models

Text Embeddings Inference (TEI) is HuggingFace's Rust server for self-hosting embedding, reranking, and sequence-classification models behind a single HTTP/gRPC API. It's for teams running RAG or search pipelines who want to own the embedding step instead of paying per-call to a hosted embeddings API.

Backend selection is hardware-aware rather than one-size-fits-all: separate Docker images per GPU compute capability (Turing, Ampere, Ada, Hopper, Blackwell) plus a candle/ONNX/Python split for architectures not yet ported to Rust. Token-based dynamic batching with an explicit max-batch-tokens knob, flash attention, and cuBLASLt are real inference-engineering levers, not just a wrapper around transformers. One server covers /embed, /rerank, /predict, and /embed_sparse, so a RAG pipeline that needs embeddings plus a cross-encoder reranker doesn't need two separate services. Per-architecture snapshot tests (backends/candle/tests/snapshots) catch numerical regressions when a new model family is added.

The GPU image matrix is a real footgun — you have to know your card's exact compute capability (75/80/86/89/90/100/120/121) and pick the matching tag, and anything below 7.5 (V100, GTX 1000s) isn't supported at all. Each server instance serves exactly one model; there's no built-in multiplexing, so running several embedding models means running and load-balancing several containers yourself. Not everything gets the 'blazing fast' Rust/candle path — some architectures still route through the Python backend, so throughput varies by model and isn't obvious until you check which backend your model actually loads. Auth is a single static bearer API key with no scoping or rate limiting, which is fine behind your own gateway but not something to expose directly to external clients.

View on GitHub → Homepage ↗

// want more like this?

We dig through GitHub every week and send a few repos picked for what you actually care about — each with an honest take like this one.

Get finds in your inbox → Search again →