finds.dev← search

// the find

alchaincyf/darwin-skill

★ 6,182 · HTML · MIT · updated Sep 2026

达尔文.skill —— 一个让你的Skill无限进化的系统:评估→改进→测试→保留或回滚 | Autoresearch-inspired autonomous skill optimization for Claude Code. Evaluate, improve, test, keep or revert.

darwin-skill is a Claude Code skill that scores an existing SKILL.md against a nine-dimension rubric, makes one targeted edit per round, re-scores the result with fresh subagent judges, and keeps the change only if the score rises. It is aimed at people who maintain many agent skills and want a repeatable, human-supervised way to tighten them.

- The keep-or-revert loop is the part worth borrowing. Each round commits before editing, changes one dimension, and rolls back with git revert instead of reset --hard, so every experiment is a separate diff you can inspect or undo.

- Judges are replaced each round and the loop stops early when a round gains less than one point. Those rules guard against the two failure modes of self-grading systems: anchoring on earlier scores and padding a skill to hit a number.

- The human checkpoints sit at phase boundaries (baseline review, each single-dimension change, regression test), so a bad edit gets caught before the next round builds on it.

- The method lives in plain Markdown instructions and a test-prompts.json file, so you can read exactly what the agent is told to do and run the skill against prompts rather than grading its text alone.

- The judge is still an LLM grading an LLM's edit. The README cites SkillLens for the claim that model self-evaluation is unreliable (46.4% accuracy) and then relies on fresh subagents from the same model family. That reduces anchoring but leaves the shared blind spot in place, and there is no published agreement check against human ratings.

- The headline numbers are anecdotal: one third-party skill going from 80.8 to 91.65 and the author's own skill going from 86.05 to 92.7. There are no repeat runs and no variance figures, so without them it is unclear whether a gap of a few points is larger than judge noise.

- It is tied to Claude Code's skill layout and subagent mechanism. references/runtime-neutrality.md suggests an intent to work elsewhere, but the loop assumes a git working tree and parallel subagents. The README asks users to commit or stash their own changes first, which is a sharp edge for anyone with uncommitted work in the skill directory.

- The academic framing is heavier than the mechanism. Most of the README's credit goes to the SkillLens and SkillOpt papers, but the concrete contribution the README describes is a nine-item rubric, a two-judge check, and an early-stop threshold. Those are reasonable heuristics, and they are not validated the way the citations suggest.

View on GitHub →

// want more like this?

We dig through GitHub every week and send a few repos picked for what you actually care about — each with an honest take like this one.

Get finds in your inbox → Search again →