finds.dev← search

// the find

egoist/sitefetch

★ 1,737 · TypeScript · MIT · updated Jan 2025

Fetch an entire site and save it as a text file (to be used with AI models).

A CLI and API for crawling a website and dumping the readable content into a single text or markdown file, aimed at feeding docs sites into LLM context windows. Useful for anyone who wants a quick offline snapshot of a docs site without hand-scraping it.

Uses Readability for content extraction instead of naive HTML stripping, so output is generally clean prose rather than nav/footer junk. Glob-based page matching (micromatch) lets you scope a crawl to a subsection of a site instead of pulling everything. Concurrency flag and a programmatic API (fetchSite) mean it's usable as a library, not just a CLI toy.

No handling described for JS-rendered sites — if content is client-side rendered, Readability has nothing to extract and you get empty pages. No rate-limiting or robots.txt respect mentioned, so cranking concurrency on someone else's site is on you to not be a jerk about it. No tests in the tree, and the README doesn't document output format specifics (chunking, size limits) which matters a lot if the whole point is feeding this into a model with a context window.

View on GitHub →

// want more like this?

We dig through GitHub every week and send a few repos picked for what you actually care about — each with an honest take like this one.

Get finds in your inbox → Search again →