// the find
fb55/htmlparser2
The fast & forgiving HTML and XML parser
htmlparser2 is a streaming HTML and XML parser with a callback interface, so it can process a document without building a tree. Its DOM, feed parsing, and stream wrappers sit on top of that core, for developers who need fast, lenient parsing in Node or the browser and can accept non-standard tree shapes.
The callback Parser never allocates a tree unless you ask for one, so streaming large documents stays cheap. The README's htmlparser-benchmark numbers (about 2.2 ms per file on real-world pages) put it at the top of that table. The tokenizer options (decodeEntities, lowerCaseTags, recognizeCDATA, recognizeSelfClosing) let you tune behavior without forking. parseFeed handles RSS, Atom, and RDF from one entry point, and the fixtures and snapshot tests in src/ back that up.
This is not spec-compliant HTML5. The README says so and points to parse5 for strict cases, but the consequence is that the tree can differ from what a browser builds, which matters for sanitizers and scrapers. The raw event API is easy to misuse: ontext can fire several times per text node, so you have to stitch the pieces together yourself. The 'one package' is really a stack of separately versioned pieces (domhandler, domutils, css-select, dom-serializer), and you need to track them together. Development is concentrated in one maintainer, which is a real bus-factor risk even though the package is widely depended on.