finds.dev← search

// the find

tjunlp-lab/Awesome-LLMs-Evaluation-Papers

★ 809 · updated May 2024

The papers are organized according to our survey: Evaluating Large Language Models: A Comprehensive Survey.

A paper list organized around a 2023 survey ('Evaluating Large Language Models: A Comprehensive Survey') that buckets LLM evaluation research into knowledge/capability, alignment, and safety categories. Useful for NLP researchers or anyone building eval harnesses who wants a map of the benchmark landscape rather than code to run.

The taxonomy mirrors the actual survey structure rather than being a flat dump, so categories like 'Multi-hop Reasoning' or 'Tool Learning' group genuinely related datasets. The badge system (Dataset/Evaluation Method/Platform/Research) lets you scan for what kind of artifact a paper produced without opening it. Each entry links both the paper and, where one exists, the GitHub repo, so it doubles as a discovery index for eval tooling, not just citations.

Stopped at the initial 2023-10-30 update despite a 2024-05-08 last push, so anything from the last 18 months of LLM eval work (which moves fast) is missing. It's a single giant README with no CONTRIBUTING.md or structured format (no JSON/YAML index), so you can't query or filter it programmatically, just ctrl-F. There's a recurring 'Rearch' typo in the badge labels throughout, and the repo has zero code or tests, so 'repo' here really just means markdown file with a nice table of contents.

View on GitHub →

// want more like this?

We dig through GitHub every week and send a few repos picked for what you actually care about — each with an honest take like this one.

Get finds in your inbox → Search again →