// the find
scrapy/scrapy
Scrapy, a fast high-level web crawling & scraping framework for Python.
Scrapy is the standard async web scraping framework for Python: you write spiders that yield requests and items, and it handles the concurrency, retries, and output formatting. It's aimed at people scraping at real scale (thousands of pages, scheduled jobs), not someone who wants to grab one page with a script.
The engine is built on an async reactor (Twisted, with asyncio support now baked in) so it runs hundreds of concurrent requests without threads or manual async plumbing. The middleware/pipeline architecture (downloader middlewares, spider middlewares, item pipelines) is genuinely extensible - things like autothrottle, robots.txt respect, cookie handling, and retry logic are all pluggable rather than hardcoded. Selectors (CSS/XPath via parsel) and built-in feed exporters (JSON/CSV/XML) mean you're not gluing lxml and csv.writer together yourself. It's been in production at Zyte-scale for over a decade, which shows in the edge-case handling (redirect loops, encoding detection, link extraction).
The callback-based request/response flow is a different mental model than requests+BeautifulSoup, and debugging a stuck spider through the Twisted/asyncio reactor is not fun for newcomers. No first-class JS rendering - anything behind client-side rendering needs scrapy-playwright or Splash bolted on, which adds its own flakiness. Settings are global and sprawling (default_settings.py is huge), so project-level config can get messy once you have more than one or two spiders with different needs. It's overkill for a one-off scrape - the project scaffolding (items.py, pipelines.py, settings.py) is wasted ceremony if you just need to pull data from fifty pages once.