// the find
different-ai/embedbase
A dead-simple API to build LLM-powered apps
Embedbase is a hosted embeddings-as-a-service API. You add text to named datasets and run semantic search over them, and a thin wrapper lets you send the results to an LLM. The repo contains the Python server with Postgres, Supabase and in-memory backends, a Next.js dashboard, and JS and Python SDKs, aimed at developers who want document retrieval without running a vector database themselves.
The storage and embedding layers are pluggable. embedbase/database/ has memory, Postgres and Supabase implementations behind one base class, and embedbase/embedding/ has OpenAI and Cohere providers behind another, so the in-memory backend lets the API run without a database. The JS SDK ships its own text splitter (src/split) with test fixtures for markdown, Python, Solidity, TypeScript and plain text. The hosted deployment is a real Docker setup, with a Dockerfile, dev and prod compose files, and an API-key auth middleware that has its own test file. CI and release workflows are separate for the server, the JS SDK and the Python SDK.
The last push was 27 November 2024, so nearly two years of changes to models, providers and dependencies have not landed. The README example still uses openai/gpt-3.5-turbo, which shows how dated the docs are. The README's own example joins search results with an empty string (.map(result => result.data).join('')), so adjacent chunks run together in the prompt, and anyone copying the snippet inherits that. The 'dead-simple API' describes the hosted service. Self-hosting means the hosted/ Docker stack, a Postgres or Supabase instance with pgvector, and OpenAI or Cohere keys, and the dashboard carries Stripe and Upstash code you would have to strip or configure. Nothing in the tree points to reranking or retrieval-quality evaluation, so the search quality is mostly taken on trust.