// the find
s0md3v/uro
declutters url lists for crawling/pentesting
uro is a Python command-line filter that takes a URL list from stdin or a file and drops URLs unlikely to add anything when crawling or testing: incremental pagination, human-written posts, static assets, and URLs that differ only in parameter values. It makes no HTTP requests, so it runs quickly on large crawl or archive dumps where the input has far more URLs than anyone can test by hand. It is aimed at bug bounty and pentest workflows.
- Deduplication keys on path and parameter names, not full URLs, so /page.php?id=1 through ?id=5000 collapse to one line with no network I/O. That is the right trade for a pre-crawl filter on large inputs.
- It is plain stdin/stdout with -i and -o for files, so it drops into a shell pipeline without wrapper code.
- The filters are explicit and composable (hasparams, noext, allexts, keepcontent, keepslash), and the built-in extension list can be overridden. When a heuristic is wrong for your target, you can turn it off rather than fork the tool.
- Collapsing by parameter names throws away value-dependent behavior. ?lang=en and ?lang=de become one URL, so a bug that only reproduces on one locale or one id range never gets tested.
- The -b option replaces the default extension list instead of extending it, so adding one extension means retyping the whole list. It is a small thing, but it is the kind of friction that gets people writing wrappers.
- Last push was 23 February 2025, about 19 months before this review. The vuln filter depends on parth, a separate project whose parameter list may have drifted, and the directory tree shows no test directory, so edge-case behavior rests on what the author checked by hand.