
rix4uni/WordList is a regularly updated collection of curated wordlists for directory fuzzing, DNS enumeration, parameter discovery and default credentials, paired with a Go generator for nuclei technology wordlists.
| Tool | rix4uni/WordList — a curated, regularly updated repository of custom wordlists for fuzzing, DNS enumeration, parameter discovery and default credentials, plus a Go generator for nuclei technology lists |
| Category | reconnaissance / wordlist corpus |
| Primary Use | Feeding content-discovery and enumeration tools such as ffuf, nuclei, and DNS bruteforcers with tailored dictionaries derived from your own authorized scope data |
| Safe Use | For penetration testers and bug bounty hunters operating strictly within authorized scopes, lab environments, and internal security assessments with written permission |
| Telemetry Note | The lists themselves are passive; the noise they generate appears in target web-server access logs and WAF events as high-rate 404/403 responses from whichever fuzzing tool consumes them |
Most content discovery fails not because the scanner is weak but because the dictionary is generic. rix4uni/WordList attacks that problem from two directions at once: it ships a corpus of curated lists covering directories, subdomains, parameters, and default credentials, and it documents a methodology for generating bespoke wordlists from the URLs and hostnames you already collected during reconnaissance of systems you are authorized to test. The repository sits at roughly 154 stars, is written primarily around Go tooling for its generator component, and carries a topic spread — bug-bounty, fuzzing, recon, osint, penetration-testing — that tells you the intended audience is working bug bounty hunters and pentest operators rather than casual users.
The philosophy of the repo is best understood from the shell pipelines embedded in its README. These are not exploit code; they are text-processing recipes that take existing recon output — a file of URLs, a file of target hostnames — and distill it into candidate words. For the custom fuzzing list fuzzing_list.txt, the pipeline strips the final path segment from each URL in urls.txt with sed, slices out path components four through eleven with cut, splits them into individual lines with tr, drops empties, and appends only new entries via anew. The result is a dictionary of path fragments that actually exist in the target's URL structure, which is far more likely to hit during directory bruteforcing than a stock list.
The DNS side follows the same pattern. The recipe for dns-wordlist.txt takes alltargets.txt, removes the trailing TLD with sed, splits on dots, filters out purely numeric labels with egrep -v '^[0-9]*$', and merges the surviving subdomain tokens through anew. This produces permutations rooted in the organization's real naming conventions — if a company prefixes hosts with api-, dev-, or internal product names, those tokens end up in your list naturally. It is a textbook example of target-informed dictionary generation, and the numeric filter is a sensible touch that removes IP-derived noise that would otherwise bloat the list.
A third pipeline prepares input for nuclei, ProjectDiscovery's template-based misconfiguration scanner. Rather than scanning every collected URL, the recipe for urls-for-nuclei.txt filters urls.txt down to URLs that actually contain a path beyond the bare origin, using grep -E "^https?://[^/]+/.+", then generates three progressively deeper truncations of each path — fields one through four, five, and six — with cut. The point is architectural: misconfiguration templates usually fire at directory level, not at deep file paths, so collapsing URLs to their first few segments deduplicates the scan surface dramatically and cuts runtime without losing template coverage.
The parameter discovery workflow is similarly economical. For params.txt, the pipeline filters urls.txt for .php? URLs, deduplicates with uro, isolates the query string after the ?, extracts the parameter name before the =, and merges with anew. PHP endpoints are targeted deliberately because they tend to expose rich legacy parameter surfaces, and a harvested parameter dictionary like this feeds directly into tools that probe for hidden parameters, forgotten debug flags, and injectable inputs — a common and legitimate bug bounty finding class.
Default credentials get their own artifact, default-username-password.txt, formatted as username:password pairs. The README shows how to split it into separate username.txt and password.txt files with cut -d":", which is the format most authentication-testing tools expect. For defenders, this file is worth knowing about precisely because default credential abuse remains one of the most trivially preventable attack vectors; if your organization deploys appliances or frameworks with vendor defaults intact, lists like this make finding them a solved problem for the attacker.
The largest single artifact is onelistforall.txt, and the README documents its provenance with unusual transparency: it is a merged, deduplicated aggregation of well-known public sources including dirsearch's dicc.txt, six2dez's OneListForAll variants, leaky-paths.txt, Bo0oM's fuzz.txt, entries from SecLists' raft-large-directories.txt, and the Assetnote httparchive-derived technology lists covering php, aspx/asp/cfm, and jsp extensions. Each source is fetched and appended through anew -q, meaning the final list is a union with duplicates stripped. Understanding the lineage matters operationally — if you already run SecLists separately, you are partly double-covering, and you should weigh list size against request budget before pointing a fuzzer at it.
The repository also imposes a consistent size taxonomy on its payload lists: *-small.txt caps at the top 50 entries, *-medium.txt at the top 500, and *-large.txt contains everything, splitting into numbered parts like *-large-1.txt and *-large-2.txt whenever a file exceeds 50 MB. This tiering is genuinely good engineering practice for fuzzing workflows. You start a broad sweep with the small variant for speed, escalate to medium on responsive targets, and reserve the large lists for deep-dive engagements where request budget and authorization scope permit sustained scanning.
The most interesting technical component is the Go program nuclei-wordlist-generator.go, stored under wordlist-generator-tools/. It generates technology-specific wordlists segmented by nuclei severity level: for each technology directory you get techname-unknown.txt, techname-info.txt, techname-low.txt, techname-medium.txt, techname-high.txt, techname-critical.txt, and a combined techname-all.txt. The implication is that the generator parses nuclei template metadata — which tags every template with a matched technology and a severity — and inverts that mapping to produce path dictionaries relevant to a detected stack. When fingerprinting says the target runs a specific CMS, you can fuzz with only the paths that template authors associate with that CMS, weighted by how severe the known issues are.
Everything here is passive from the tooling side: the lists are inert text, the pipelines run on your own recon data, and the only network interaction is fetching raw files from GitHub and public CDNs during list construction. The operational footprint lands wherever you point your consuming tools — ffuf, nuclei, or a DNS bruteforcer — and manifests as elevated 404 rates in target access logs and WAF telemetry. On unauthorized targets that is both illegal and immediately visible; within a contracted scope or a lab it is exactly the observability your blue team counterpart should be using to validate detection tuning.
Getting the corpus is a single git clone https://github.com/rix4uni/WordList into your recon tooling directory, after which the flat text files can be passed to any tool accepting a -w wordlist flag. The README pipelines additionally depend on anew, uro, and standard coreutils, all small and widely available. One caveat worth noting: the repository declares no license, which strictly speaking means no redistribution rights are granted — treat it as a personal-use resource rather than something to vendor into a commercial product without contacting the author.
As a benchmark for what curated recon data looks like in 2026, rix4uni/WordList is a solid reference point. It will not replace SecLists as the canonical corpus, but the combination of severity-segmented technology wordlists, size-tiered payload variants, and documented pipelines for deriving dictionaries from your own collected URLs makes it more of a methodology artifact than a simple list dump — and that methodology, applied to your authorized scope, is where the real value sits.
rix4uni/WordList.Educational analysis for authorized security professionals. Use only in controlled, authorized environments.
0 comentários:
Post a Comment
Note: Only a member of this blog may post a comment.