finds.dev← search

// the find

agentscope-ai/OpenJudge

★ 852 · Python · Apache-2.0 · updated Sep 2026

OpenJudge: A Unified Framework for Holistic Evaluation and Quality Rewards

OpenJudge is an evaluation framework for AI agents and LLM apps: a library of 50+ graders (relevance, hallucination, tool selection, trajectory quality, etc.) plus tooling to generate custom rubrics and train dedicated judge models. It's aimed at teams building eval pipelines for agents or wanting to turn grading output into RLHF-style reward signals.

The grader taxonomy actually covers the agent lifecycle (tool calls, memory, plan feasibility, trajectory) instead of just scoring final text output, which is where most eval libraries stop. It offers a real progression from zero-shot rubric generation (no labeled data) to data-driven rubric generation (from labeled examples) to full judge-model training (SFT, GRPO pointwise/pairwise, Bradley-Terry), with runnable scripts in cookbooks/training_judge_model rather than just prose. Integrations with LangSmith, Langfuse, and VERL mean it can plug into an existing observability or RL-training stack instead of being a silo.

v0.2.0 is a hard breaking change from the old rm-gallery package (different PyPI name, different import namespace) — anyone on v0.1.x has a real migration, not a version bump. Grader trustworthiness is entirely downstream of whatever judge LLM you plug in (examples default to Qwen); the framework gives you the harness but not a strong default judge, so you still have to run the validation step yourself. The training pipelines pull in a heavy RL stack (verl) just to exercise the reward-function code, and a lot of substantive documentation lives off-repo on the docs site and openjudge.me rather than in the README/cookbooks themselves. It's also moving fast (skills graders, PawBench, new leaderboards all shipped within months of the v0.2 rename), so expect rubric/config formats to keep shifting.

View on GitHub → Homepage ↗

// want more like this?

We dig through GitHub every week and send a few repos picked for what you actually care about — each with an honest take like this one.

Get finds in your inbox → Search again →