// the find
mathiasbynens/he
A robust HTML entity encoder/decoder written in JavaScript.
he is a spec-compliant HTML entity encoder/decoder for JavaScript, built to replicate the exact tokenization algorithm from the HTML5 spec rather than approximate it with regex. It's for anyone who needs entity handling to be actually correct — HTML parsers, sanitizers, linters, static site generators — not just for turning `&` into `&`.
Implements the real WHATWG tokenizer state machine for character references, including ambiguous ampersand handling, which most competing libraries get wrong on edge cases. Correctly handles astral plane symbols (surrogate pairs) where regex-based encoders silently corrupt them. Zero runtime dependencies, and the data tables (legacy named refs, decode maps) are generated from the spec itself rather than hand-maintained, so behavior is verifiable against the source of truth. Ships both a programmatic API and a CLI binary for pipeline use.
No commits since December 2021 — for a package this small in scope that's arguably fine, but there's been zero response to any issues or PRs in ~4 years, so don't expect fixes if you hit something. CommonJS/UMD only, no ESM build, so it doesn't tree-shake and adds friction in modern bundler setups. No TypeScript types in the package itself; you're pulling `@types/he` separately and hoping it's still in sync. The data JSON files are baked into the published package, so you're shipping the full named-character-reference table even if you only ever decode numeric entities.