// the find
agentscope-ai/OpenJudge
OpenJudge: A Unified Framework for Holistic Evaluation and Quality Rewards
OpenJudge is an evaluation framework for AI agents and LLM apps: a library of 50+ graders (relevance, hallucination, tool selection, trajectory quality, etc.) plus tooling to generate custom rubrics and train dedicated judge models. It's aimed at teams building eval pipelines for agents or wanting to turn grading output into RLHF-style reward signals.
The grader taxonomy actually covers the agent lifecycle (tool calls, memory, plan feasibility, trajectory) instead of just scoring final text output, which is where most eval libraries stop. It offers a real progression from zero-shot rubric generation (no labeled data) to data-driven rubric generation (from labeled examples) to full judge-model training (SFT, GRPO pointwise/pairwise, Bradley-Terry), with runnable scripts in cookbooks/training_judge_model rather than just prose. Integrations with LangSmith, Langfuse, and VERL mean it can plug into an existing observability or RL-training stack instead of being a silo.
v0.2.0 is a hard breaking change from the old rm-gallery package (different PyPI name, different import namespace) — anyone on v0.1.x has a real migration, not a version bump. Grader trustworthiness is entirely downstream of whatever judge LLM you plug in (examples default to Qwen); the framework gives you the harness but not a strong default judge, so you still have to run the validation step yourself. The training pipelines pull in a heavy RL stack (verl) just to exercise the reward-function code, and a lot of substantive documentation lives off-repo on the docs site and openjudge.me rather than in the README/cookbooks themselves. It's also moving fast (skills graders, PawBench, new leaderboards all shipped within months of the v0.2 rename), so expect rubric/config formats to keep shifting.