finds.dev← search

// the find

danielmiessler/RobotsDisallowed

★ 1,494 · Shell · updated Aug 2022

A curated list of the most common and most interesting robots.txt disallowed directories.

RobotsDisallowed is a set of static wordlists scraped from the robots.txt Disallow entries of the Alexa/Majestic top 100K sites, meant to seed content-discovery during web security assessments and bug bounty recon. It's for pentesters and bug bounty hunters who want a pre-built directory wordlist instead of building one from scratch.

curated.txt distills the raw dump to ~500 entries matching sensitive-sounding strings (admin, login, backup, etc.) plus the top 25 most common paths — a genuinely useful signal-to-noise cut over the raw list. The archive/code directory keeps the original scraping and cleanup shell scripts, so the methodology is transparent rather than a black-box wordlist. Multiple list sizes (top10 through top100000) let you trade scan time against coverage depending on assessment budget.

The underlying data is from March 2019 — over six years stale. Site structures and CMS defaults have shifted a lot since then, so a good chunk of these paths are dead weight now. There's no CI or scheduled job to re-scrape, and regenerating it manually is harder than it should be since Alexa (one of the two original sources) doesn't exist anymore. It's just text files and loose shell scripts with no packaging or dependency management, so integrating it into a scan pipeline means manual copy-paste. Effectively unmaintained: last real content update 2019, last push 2022, no meaningful issue or PR activity since.

View on GitHub →

// want more like this?

We dig through GitHub every week and send a few repos picked for what you actually care about — each with an honest take like this one.

Get finds in your inbox → Search again →