
findcdn is a Python tool from cisagov that fingerprints the Content Delivery Network behind any domain by scraping HTTPS headers, CNAME records, and WHOIS data for reconnaissance and defensive asset mapping.
| Tool | cisagov/findcdn — CDN detection tool that scans domains and identifies which Content Distribution Network serves them |
| Category | OSINT / domain reconnaissance (Python) |
| Primary Use | Bulk-enumerating domains to determine which CDN each uses, output as JSON, from the CLI (findcdn list) or as an importable Python module |
| Safe Use | Mapping CDN coverage across your own or client-authorized asset inventories, attack-surface reviews, and defensive research on infrastructure composition — it performs passive lookups and header checks, not exploitation |
| Telemetry Note | Generates ordinary HTTPS requests, CNAME/DNS lookups, and WHOIS queries against targets; the default user_agent is overridable via --user_agent, so defenders may see bursts of same-UA requests resolved to known CDN edges |
findcdn is a Python 3.7+ utility maintained by the Cybersecurity and Infrastructure Security Agency (cisagov) that answers a deceptively simple question at scale: which Content Distribution Network is a given domain actually using? The tool accepts domains either inline via findcdn list or from a file via findcdn file, processes them across a configurable thread pool, and emits a structured JSON result mapping each domain to its detected CDN. That dual identity — a standalone CLI and an importable library — is deliberate, and the README treats both paths as first-class citizens rather than bolting the module interface on afterward.
The detection engine is the interesting part of the architecture. Rather than relying on a single signal, findcdn triangulates across three independent data sources: HTTPS server headers, CNAME DNS records, and WHOIS data. Each source is fingerprinted against a curated signature list defined in cdn_config.py (found under src/findcdn/cdnEngine/detectCDN/), which is effectively the extensible knowledge base of the whole project. Adding support for a new CDN means extending that config rather than touching engine code — a clean separation that makes the repository approachable for contributors.
Internally the README describes a three-stage pipeline. A main runner validates and organizes the input domains, then orchestrates the CDN Engine. The engine stages domains in a pot and hands them to a component the authors call Chef, which invokes the detection library and then performs the analysis that sets a boolean has_cdn value per domain before returning results upstream. Finally, CDN Detection performs the actual scraping and fingerprint matching. The whimsical naming (pot, Chef) shouldn't obscure that this is a reasonable orchestration design: parse, batch, detect, aggregate, serialize.
Installation is standard pip fare, either from the repository's requirements.txt or directly with pip install git+https://github.com/cisagov/findcdn.git. The authors recommend a virtual environment managed through pyenv and pyenv-virtualenv, which is a small detail but signals the government-origin code hygiene the repo carries elsewhere: build badges, Coveralls coverage reporting, historical LGTM alerts, and Snyk vulnerability scanning are all wired into the README. This is maintained code with CI, not a throwaway script.
The CLI surface is compact but well-considered. Beyond -o/--output=FILE for writing the JSON to disk and -v for verbosity, the operationally significant flags are -t/--threads=<thread_count> for concurrency control, --timeout to bound how long any single domain can stall the run, --user_agent to control the client fingerprint presented to targets, --all to include CDN-less domains in output rather than filtering them, and -d/--double, which runs the checks twice per domain to increase accuracy. The timeout and thread controls matter when you are feeding thousands of domains through the tool and a handful of blackholed or slow endpoints would otherwise dominate wall-clock time.
Sample output in the README shows the shape of the result: a JSON object containing a date, a CDN_count, and a domains dictionary where each entry carries the resolved IP, the raw CDN suffix match (for example .cloudflare.com), and a normalized cdns_by_names label (for example Cloudflare). The example run — findcdn list asu.edu -t 7 --double — completes in about a second with a two-thread progress bar, so single-domain lookups are effectively instant while bulk scans scale with thread count.
The library interface mirrors the CLI almost one-to-one, exposed as findcdn.main() with keyword arguments: domain_list, output_path, verbose, all_domains, interactive (the progress bar), double_in, threads, timeout, and user_agent. The README's example imports findcdn, passes a list of domains with output_path="output.json" and double_in=True, then parses the returned JSON with the standard json module to iterate dumped_json['domains']. This makes the tool trivially embeddable in larger asset-management or reconnaissance pipelines where CDN attribution is one enrichment step among many.
The most candid section of the README is the History. findcdn originally aimed to automatically determine whether a CDN-fronted domain was domain-frontable — that is, usable for domain fronting, the technique of disguising traffic to one hostname as traffic to another on the same CDN. The authors abandoned that goal because reliable frontability detection across every CDN provider carried significant overhead, and pivoted the project to pure CDN *detection*. The repository wiki retains their research notes, design decisions, and fronting playbooks, and contributions of additional frontable domains are explicitly invited.
That history is worth understanding from a defensive standpoint as much as an offensive one. Knowing which CDN an organization's domains ride on is directly relevant to defenders assessing exposure to CDN-side misconfiguration, origin bypass, and fronting abuse — if an adversary can front through a CDN you also use, your egress allow-listing and domain reputation controls may behave unexpectedly. findcdn gives you the inventory layer for that conversation: cheap, repeatable, structured answers about CDN composition across an entire domain set.
In an authorized workflow, the natural fit is early recon or asset triage: feed it a harvested domain list, get back JSON annotated with CDN attribution, and pivot downstream with tooling appropriate to your scope. Because the tool's footprint is limited to ordinary HTTPS requests, DNS lookups, and WHOIS queries, it sits comfortably within the passive-to-light-active end of the spectrum and produces no exploitation traffic. The --user_agent flag lets operators match their tooling profile to their engagement conventions, and --double trades runtime for confidence where a single-pass fingerprint might be ambiguous.
Caveats worth noting: the supported CDN list lives entirely in cdn_config.py, so coverage quality depends on how current that file is — smaller or regional CDNs may simply not be in the signature set, and detection fidelity degrades for providers that aggressively strip identifying headers. The Python 3.7+ requirement and the absence of Python 2 support are stated plainly, and the project is released under CC0 1.0 as worldwide public domain, which removes any licensing friction for embedding it in government, commercial, or personal tooling alike.
All told, findcdn is a focused, well-instrumented utility that does one thing cleanly: CDN attribution at domain-list scale, with honest documentation about what it is not (a frontability oracle) and a clean extension point in cdn_config.py for what it is. For professionals building asset inventories or researching CDN ecosystem composition, it earns its place as a dependable enrichment primitive rather than a headline tool.
cisagov/findcdn.Educational analysis for authorized security professionals. Use only in controlled, authorized environments.
0 comentários:
Post a Comment
Note: Only a member of this blog may post a comment.