// the find
cheeriojs/cheerio
The fast, flexible, and elegant library for parsing and manipulating HTML and XML.
Cheerio is a server-side library for parsing and editing HTML and XML with a jQuery-style API, running on Node, Bun, Deno, and in the browser. It suits scraping and markup transformation with CSS selectors when you don't need a real DOM, script execution, or layout.
- The default parser is parse5, which follows the WHATWG HTML spec, so malformed markup gets repaired the way a browser would repair it. The optional htmlparser2 backend and the slim entry point (src/slim.ts) trade some of that fidelity for speed and more forgiving XML handling.
- The TypeScript types live in the source tree, not in a separate @types package, and the selection and traversal API is small enough to learn from the README.
- The maintenance setup is more than most libraries this size bother with: CI, CodeQL, zizmor, OpenSSF Scorecard, a threat model, a security policy, and a benchmark workflow that measures the speed claim instead of just asserting it.
- There is no layout, no computed styles, and no script execution. Client-rendered pages need a headless browser before cheerio sees any content, and the README does not make that clear up front.
- It implements a subset of jQuery, and the 'DOM Node object' section is a loose approximation of browser nodes. Code ported from jQuery or from browser DOM code will hit missing methods and property differences, so check the API docs before assuming parity.
- The README leads with 'blazingly fast' but publishes no numbers. The benchmark sits in a separate folder, so anyone choosing between parse5 and htmlparser2 has to measure on their own documents.