finds.dev← search

// the find

jeinlee1991/chinese-llm-benchmark

★ 6,307 · updated Jul 2026

非线智能 NoneLinear - ReLE评测:中文AI大模型能力评测(持续更新):目前已囊括374个大模型,覆盖chatgpt、gpt-5.4、谷歌gemini-3.1-pro、Claude-4.6、文心ERNIE-X1.1、ERNIE-5.0、qwen3.6-max、qwen3.6-plus、百川、讯飞星火、商汤senseChat等商用模型, 以及step3.5-flash、kimi-k2.6、ernie4.5、MiniMax-M2.7、deepseek-v4、Qwen3.6、llama4、智谱GLM-5.1、MiMo-V2、LongCat、gemma4、mistral等开源大模型。不仅提供排行榜,也提供规模超200万的大模型缺陷库!方便广大社区研究分析、改进大模型。

A Chinese-language LLM benchmark that has been running since 2023 and now covers 394 models across 7 domains and ~300 subcategories, from primary school math to bar exams to medical licensing. It's aimed at Chinese developers and researchers who need to compare models on tasks that actually matter in the Chinese market, not just MMLU reruns. The 2M+ bad-case database is the most distinctive asset — it lets you see exactly where a model fails, not just its aggregate score.

The bad-case library is genuinely useful — for each benchmark you can browse actual model failures, which is far more actionable than a leaderboard number. Coverage of Chinese professional exams (medical licensing, bar exam, CPA, civil service) is deep and specific in a way that Western benchmarks simply don't attempt. Update cadence is real: the changelog shows weekly additions of new models, and the 2025 gaokao questions were added days after the exam. The custom filtering tool at nonelinear.com lets you build your own ranked slice across any dimension combination, which saves a lot of manual cross-referencing.

There is no code in this repo — it's entirely data and markdown, so you can't reproduce the evaluation pipeline or verify methodology independently. The scoring formula (0.3 professional + 0.7 general, average of averages) buries a lot of important choices that aren't explained or justified. The repo is effectively a marketing vehicle for NoneLinear's commercial API gateway, which is prominently plugged mid-README and makes it hard to tell where the benchmark ends and the upsell begins. Most of the benchmark data is multiple-choice on Chinese professional exams, which measures knowledge recall well but says little about generation quality, reasoning under ambiguity, or real-world task completion.

View on GitHub → Homepage ↗

// want more like this?

We dig through GitHub every week and send a few repos picked for what you actually care about — each with an honest take like this one.

Get finds in your inbox → Search again →