finds.dev← search

// the find

mlabonne/llm-datasets

★ 4,801 · updated Apr 2026

Curated list of datasets and tools for post-training.

A README-only awesome-list of instruction, preference, and reasoning datasets for LLM post-training (SFT/DPO), maintained by Maxime Labonne. Aimed at ML engineers assembling a fine-tuning data mix who don't want to start from a blank HuggingFace search.

Organized by actual use case (general, math, code, instruction-following, multilingual, agent/function-calling, real conversations, preference) instead of a flat link dump. Each row carries sample count, a thinking/reasoning-trace flag, and license notes where it matters, which saves a click into the dataset card. Entries are dated and current through April 2026, so it's tracking the newest large releases (Nemotron Post-Training v3, Dolci, SYNTHETIC-2) rather than going stale. Pairs the dataset tables with a tools section (scraping, filtering, generation, exploration), so it also doubles as a pipeline reference, not just a bibliography.

It's a static README, not a package or API — nothing to install, nothing to query, you're clicking through to HuggingFace for everything. License info is inconsistent: some rows flag CC-BY-NC or NVIDIA's custom license, most say nothing, so you still have to check before using anything commercially. No guidance on which overlapping general-purpose mixture to actually pick for a given model size or compute budget — several entries compete for the same slot with no tiebreaker. Quality signal is just the maintainer's one-line note; there's no dedup or overlap check across datasets, so blending a few of these blind risks duplicate samples.

View on GitHub → Homepage ↗

// want more like this?

We dig through GitHub every week and send a few repos picked for what you actually care about — each with an honest take like this one.

Get finds in your inbox → Search again →