finds.dev← search

// the find

mckaywrigley/clarity-ai

★ 1,422 · TypeScript · MIT · updated Mar 2024

A Perplexity clone.

Clarity AI is a minimal Perplexity clone built on Next.js pages router: it scrapes Google for a query, pulls text from the resulting pages, stuffs it into a prompt, and streams an OpenAI-generated answer back to the user. It's aimed at developers who want to see the whole RAG-over-search pattern in a handful of files, not at anyone looking for a production search product.

The entire pipeline (scrape -> parse -> prompt -> stream) fits in about half a dozen files under pages/api and utils, so you can read the whole thing in one sitting and actually understand what's happening at each step. Streaming the OpenAI response back to the client is wired up correctly, which is the part people usually get wrong on a first pass. It has no unnecessary abstraction layers, state management libraries, or config system getting in the way of the core idea.

Scraping Google directly for sources is fragile and will break or get rate-limited/blocked outside light personal use; the README itself flags this as a known problem it never fixed. It's hardcoded to OpenAI's now-deprecated text-davinci-003 for source handling, so getting it working on current chat models requires real surgery, not a config change. There's no error handling around the scrape step, no tests, and no rate limiting or caching, so a few concurrent users would either blow through the OpenAI key's budget or get IP-blocked by Google. The project has been dormant since March 2024 and was explicitly published as a demo/teaching repo, not something intended to be maintained or extended.

View on GitHub → Homepage ↗

// want more like this?

We dig through GitHub every week and send a few repos picked for what you actually care about — each with an honest take like this one.

Get finds in your inbox → Search again →