// the find
unclecode/crawl4ai
Open-source web crawler and scraper for LLMs and AI agents: any website into clean, LLM-ready Markdown. Run it yourself, or use Crawl4AI Cloud with one key.
A Python library and Docker server that drives a Playwright browser to turn web pages into Markdown for RAG pipelines and agents, with CSS, XPath and LLM-based structured extraction on top. It is for developers building ingestion pipelines who want the crawler under their own control rather than a hosted API. The README spends a lot of its length on the paid cloud product, so read it for the library, not the pitch.
The Markdown path is the part worth reading. PruningContentFilterLXML and BM25ContentFilter produce fit_markdown next to raw_markdown, so you can see what the filter removed instead of trusting it blind. Most extraction does not need an LLM: JsonCssExtractionStrategy and the XPath variant handle repeated-structure pages cheaply, and LLMExtractionStrategy goes through LiteLLM, so the provider is swappable. Deep crawl is a real feature set, with BFS, DFS and best-first strategies, resume_state for crash recovery, and arun_many on a memory-adaptive dispatcher. The Docker server requires a bearer token on every endpoint by default, which most self-hosted scrapers do not bother with.
The README is a sales page for the hosted product. Cloud badges, a funding appeal, sponsor tables and a 'Which one?' comparison sit above the install steps, and the search and answer endpoints are cloud-only. It is fair to ask how the roadmap is weighted between the library and the service. Bot walls and JavaScript-heavy sites are the hard part of crawling, and the README sells their handling as a cloud feature ('handled for you, automatically'). Self-hosted, you bring your own proxies and tune the stealth and browser settings yourself, and there is no benchmark against current bot detection. The README says you must include attribution, which goes further than Apache 2.0 requires, so read LICENSE and any NOTICE file before deciding it does not apply to you. Browsers are a hard dependency: crawl4ai-setup, and a manual playwright install --with-deps when that fails. Expect that to break first in CI and in slim containers.