finds.dev← search

// the find

guillaumemeyer/watermarks-remover

★ 23,331 · Python · MIT · updated Oct 2026

A privacy-first app that strips AI watermarks from content you own.

This is a stdlib-only Python HTTP service and thin agent skill that strips provenance marks from files (C2PA manifests, EXIF, XMP, document properties) and runs a rewrite pass over text to weaken statistical watermarks. It is aimed at people cleaning content they own before publishing it. The file-metadata side has the clearest mechanics; the text side depends on a model backend you configure yourself.

- The core is stdlib-only Python 3.10+ with no install step. External tools (c2patool, exiftool, qpdf) are version-probed rather than trusted on PATH, and /capabilities reports only what actually runs.

- Format routing is by magic bytes, and anything unrecognized is labelled unknown and refused by auto-clean. The text tools now refuse binary input after an earlier version decoded DOCX bytes as text and wrote them back mangled. The README documents that failure and the fix, which is more useful than a changelog line.

- The check and clean paths share one scanner (audit_lib), so the PostToolUse hook, the pre-commit gate, and the SARIF export agree on what counts as actionable. The hook defaults to report-only, and clean mode swaps files only on a real difference so unchanged files keep their mtimes.

- Detectors are fail-soft. An unconfigured, timed-out, or erroring detector reports available: false and never blocks the clean, so a missing GPU sidecar cannot take down the whole pipeline.

- Text cleaning does not work out of the box. POST /clean on text returns 400 unless a Layer B backend is configured, either transformers with roberta-large or an LLM endpoint for the paraphrase step. Adopting the text side means adopting a model dependency. Without a detector configured, the only automatic quality signal is bigram-Jaccard divergence, which measures how different the output is, not whether it still says the same thing.

- Nothing here can confirm a removal against the real vendor detectors. MarkLLM and the keyed-Gumbel replay only verify same-config marks, and the Claude text detector is a placeholder. The README says this plainly, but the Layer A/B table reads as if the outcome is settled.

- The vendor coverage is partly historical. The README says Gemini SynthID text support was removed in August 2026 when Google retired it on the API, so the vendor table is partly a record of what used to work.

- The heavy image stack has licensing the install path does not surface. The reverse-SynthID scorer is under a non-commercial research license and noai-watermark ships no LICENSE, so both stay local-build only. The compose profile hides that behind one command, and the CtrlRegen venv pins packages the README itself says carry published advisories.

View on GitHub → Homepage ↗

// want more like this?

We dig through GitHub every week and send a few repos picked for what you actually care about — each with an honest take like this one.

Get finds in your inbox → Search again →