// the find
egoist/sitefetch
Fetch an entire site and save it as a text file (to be used with AI models).
A CLI and API for crawling a website and dumping the readable content into a single text or markdown file, aimed at feeding docs sites into LLM context windows. Useful for anyone who wants a quick offline snapshot of a docs site without hand-scraping it.
Uses Readability for content extraction instead of naive HTML stripping, so output is generally clean prose rather than nav/footer junk. Glob-based page matching (micromatch) lets you scope a crawl to a subsection of a site instead of pulling everything. Concurrency flag and a programmatic API (fetchSite) mean it's usable as a library, not just a CLI toy.
No handling described for JS-rendered sites — if content is client-side rendered, Readability has nothing to extract and you get empty pages. No rate-limiting or robots.txt respect mentioned, so cranking concurrency on someone else's site is on you to not be a jerk about it. No tests in the tree, and the README doesn't document output format specifics (chunking, size limits) which matters a lot if the whole point is feeding this into a model with a context window.