// the find
JuliusBrussee/caveman
🪨 why use many token when few token do trick. Viral skill + proxy for coding agents that cuts 65% of tokens by talking like a caveman.
Caveman is a set of agent skills plus a local CLI proxy. The skill makes a coding agent answer in terse fragments while leaving code, commands, and error text untouched, and the proxy compresses logs, JSON, CSV, YAML, and test output before the agent reads them. It is for developers who already pay real token bills in Claude Code or a similar agent and can live with a clipped voice.
The proxy keeps every original on disk and lets the agent fetch it back, so compression is recoverable instead of silently lossy. The benchmarks come with their method, per-file breakdowns, and a HONEST-NUMBERS doc that lists where the tool loses; the HTML row shows a 9.9% session regression rather than hiding it. The skill draws a sensible line for an agent that writes code: articles can go, but negations, numbers, paths, and payloads stay verbatim, which is where terse prompting usually breaks. Targeting what the agent reads, not only what it says, is the more useful half of the idea.
The description says it cuts 65% of tokens, but the README's own measured figures are 33.2% fewer input tokens across whole proxy sessions and 8.5% fewer output tokens in the JetBrains A/B, so anyone quoting the headline should use the lower numbers. The rules add about 1,000 input tokens to every call, so short exchanges can come out net worse, and HTML gets no compressor yet. The CLI sends usage telemetry by default (commands run, token counts, install ID, OS, and IP). It is opt-out, which is worth knowing before it lands on a work machine. The repo also spans an npm CLI, a Go engine, a Go browse tool, TypeScript and Python middleware, and 30+ agent profiles, so it is hard to tell which parts are maintained core and which are experiments.