// the find
huggingface/evaluation-guidebook
Sharing both practical insights and theoretical knowledge about LLM evaluation that we gathered while managing the Open LLM Leaderboard and designing lighteval!
A long-form, free-text guide to LLM evaluation written by the Hugging Face team behind the Open LLM Leaderboard and lighteval, covering automated benchmarks, human eval, and LLM-as-judge approaches. It's aimed at anyone building eval pipelines, from beginners who need the conceptual basics to practitioners who want the troubleshooting and tips-and-tricks sections.
The troubleshooting chapters (inference, reproducibility, math parsing) are the real value here — they document specific failure modes like tokenizer mismatches and non-determinism that you'd otherwise only learn by getting burned. It's organized by skill level (basics vs. tips-and-tricks vs. designing-your-own) so you can skip straight to what's useful instead of reading linearly. Content comes from people who actually ran a leaderboard at scale, so the advice reflects real operational pain rather than textbook theory.
The README itself says this version is no longer maintained and points to a HF Spaces page as the current source — so you're looking at a frozen snapshot, and anyone finding this via GitHub search is one click from a dead end. It's pure markdown/prose with no code, scripts, or runnable notebooks despite being tagged Jupyter Notebook as the primary language, so there's nothing here to actually execute or test against your own eval setup. Licensed CC-BY-NC-SA, which blocks commercial reuse of the content itself if you wanted to adapt it into internal docs.