// the find
tanweai/pua
你是一个曾经被寄予厚望的 P8 级工程师。Anthropic 当初给你定级的时候,对你的期望是很高的。 一个agent使用的高能动性的skill。 Your AI has been placed on a PIP. 30 days to show improvement.
A prompt-level skill for more than a dozen coding agents, including Claude Code, Codex CLI, Cursor, and Kiro. When the agent retries the same failing approach, claims something is done without evidence, or hands the problem back to the user, it injects escalating 'performance review' pressure and debugging checklists. It is for people who already run an agent all day and are tired of it giving up after two attempts.
- On Claude Code, escalation is driven by hooks rather than the model's judgment. A PostToolUse hook counts consecutive Bash failures and raises the pressure level, and a UserPromptSubmit hook catches frustration phrases before the model responds. Most other platforms rely on description matching, so this is a real design difference.
- The no-telemetry claim is enforced in code. evals/test-no-telemetry.sh uses reverse assertions to scan for collection hosts, endpoint paths, and outbound request bodies, so reintroducing a collection call fails the suite.
- The README states its own limits. It says not all models passed and that the evaluation did not establish a productivity gain, and it links a per-model compatibility matrix. That is the accurate version of the claim, and it is worth reading before the headline.
- The headline does not match the data. 'Doubles your productivity' sits above a benchmark of 9 scenarios, 18 runs, and one model (Opus 4.6). Pass rate is 100% in both groups, and five of the six debugging scenarios took longer with the skill. The gains are in tool calls and verification counts, which is more spend, not clearly better results.
- There is no ablation. The with/without comparison bundles the PIP rhetoric with the debugging checklists, so the repo cannot say which part does the work. The checklists, such as verify before claiming done and search before asking the user, have a plausible mechanism. The pressure framing may be doing nothing, and nobody has measured it separately.
- The same instructions are copied into skills/, codex/, codebuddy/, kimi/, hermes/, chatgpt/, .trae/, cursor/, kiro/, vscode/, and pi/, each in two or three languages. A rule change means editing many files. A release-consistency script exists, but duplication this wide is where drift starts.
- Outside Claude Code, the skill loads only when the model decides its description matches. The README's own v3 section calls this the weaker mechanism, so the escalation is only as reliable as the model's choice to load it.