// the find
rashadphz/farfalle
🔍 AI search engine - self-host with local or cloud LLMs
A self-hosted Perplexity clone: a FastAPI backend that fans a query out to a search provider and an LLM, then streams back a synthesized answer with citations. For developers who want a private search assistant they can point at local models via Ollama instead of paying for someone else's.
The search backend is abstracted behind a provider interface (search/providers/base.py) with Tavily, Serper, Bing, and SearXNG as drop-in implementations, so you're not locked into a paid API just to get results. Same deal on the LLM side — Ollama for local models, OpenAI/Groq for cloud, and LiteLLM as an escape hatch for anything else, all through one base.py abstraction. It ships with actual infra concerns handled rather than bolted on later: Alembic migrations for Postgres, Redis-backed rate limiting, and Logfire for structured logging. The docker-compose.dev.yaml gets you a running stack in one command, which is more than most repos in this space bother with.
No tests anywhere in the tree — for a service juggling multiple LLM and search provider integrations, that's a real risk when any one of those upstream APIs changes shape. Last push was over a year ago and the roadmap's last unchecked item (chat with local files) never landed, so this reads as abandoned rather than paused. Getting the full feature set working means collecting up to five separate API keys (Tavily, Serper, Bing, OpenAI, Groq) or standing up SearXNG yourself — the 'self-host for free' pitch still has real setup cost. The generated TypeScript API client (src/frontend/generated/) is committed to the repo, which means backend schema changes require remembering to regenerate and commit it rather than it happening automatically in CI.