// the find
huggingface/text-embeddings-inference
A blazing fast inference solution for text embeddings models
Text Embeddings Inference (TEI) is HuggingFace's Rust server for self-hosting embedding, reranking, and sequence-classification models behind a single HTTP/gRPC API. It's for teams running RAG or search pipelines who want to own the embedding step instead of paying per-call to a hosted embeddings API.
Backend selection is hardware-aware rather than one-size-fits-all: separate Docker images per GPU compute capability (Turing, Ampere, Ada, Hopper, Blackwell) plus a candle/ONNX/Python split for architectures not yet ported to Rust. Token-based dynamic batching with an explicit max-batch-tokens knob, flash attention, and cuBLASLt are real inference-engineering levers, not just a wrapper around transformers. One server covers /embed, /rerank, /predict, and /embed_sparse, so a RAG pipeline that needs embeddings plus a cross-encoder reranker doesn't need two separate services. Per-architecture snapshot tests (backends/candle/tests/snapshots) catch numerical regressions when a new model family is added.
The GPU image matrix is a real footgun — you have to know your card's exact compute capability (75/80/86/89/90/100/120/121) and pick the matching tag, and anything below 7.5 (V100, GTX 1000s) isn't supported at all. Each server instance serves exactly one model; there's no built-in multiplexing, so running several embedding models means running and load-balancing several containers yourself. Not everything gets the 'blazing fast' Rust/candle path — some architectures still route through the Python backend, so throughput varies by model and isn't obvious until you check which backend your model actually loads. Auth is a single static bearer API key with no scoping or rate limiting, which is fine behind your own gateway but not something to expose directly to external clients.