// the find
mlabonne/llm-autoeval
Automatically evaluate your LLMs in Google Colab
A Colab notebook plus a shell script that spins up a RunPod GPU pod, runs one of three benchmark suites (Nous, lighteval, or Open LLM) against a Hugging Face model, and dumps the results to a GitHub Gist. Built by Maxime Labonne mainly to feed his own 'Yet Another LLM Leaderboard' space, useful for anyone who wants a quick, repeatable eval run without hand-configuring lm-evaluation-harness or lighteval themselves.
Wraps three different eval backends (Nous via lm-eval-harness, lighteval, Open LLM via vllm) behind one consistent interface, so you don't have to learn each library's CLI separately. The Gist upload step means every run produces a shareable, comparable artifact instead of a local JSON file nobody looks at again. Using vllm for the Open LLM suite is a real speed win over the harness's native backend.
The README states outright it's 'in the early stages and primarily designed for personal use' — this isn't hardened tooling, it's one person's eval harness made public. Last push was May 2024, so it predates a lot of churn in lighteval and the HF eval ecosystem; expect version skew. It's hard-wired to RunPod (a paid service you have to fund manually) and Colab secrets, with no local/offline execution path. The troubleshooting section is basically a list of unresolved issues — 'outdated CUDA drivers: that's unlucky, start a new pod' — rather than fixes baked into the script, and mmlu is flagged as missing from the Open LLM suite due to an unresolved vllm problem.