
Exploit-Index aggregates roughly 150,000 verified PoC and exploit references across ~250,000 CVEs into a zero-API JSON dataset, aimed at authorized SOC, DevSecOps, and threat-intelligence workflows.
| Tool | SecureWithUmer/Exploit-Index — daily-updated, zero-API index correlating CVEs with PoCs, Exploit-DB entries, Nuclei templates, and Metasploit modules |
| Category | Vulnerability intelligence aggregation / CVE enrichment dataset |
| Primary Use | Local lookups of exploit availability, EPSS, CVSS signals, and CISA KEV status to prioritize patching and triage during authorized assessments |
| Safe Use | Intended for authorized security teams, SOC analysts, and researchers performing defensive triage and vulnerability prioritization on systems they are licensed to assess |
| Telemetry Note | Purely a passive data source; it makes no network contact at query time beyond git clone/curl pulls from GitHub raw content, so its use is essentially invisible to targets |
Exploit-Index positions itself as a single aggregated reference layer over the fragmented world of public vulnerability data. Rather than being an exploitation tool, it is a correlation engine and dataset: it normalizes roughly 250,000 CVE records and claims more than 150,000 verified exploit or PoC references into one canonical JSON index that can be queried locally with no API keys, no tokens, and no rate limits. The repository currently sits at 64 stars, is written in Python, and is released under an MIT license with its default branch on main. For teams that routinely bounce between the NVD, Exploit-DB, scattered GitHub PoC repos, and the CISA KEV catalog, the pitch is consolidation: one fetch, one file, one jq filter.
The architecture, as described in the README, is a fault-tolerant GitHub Actions pipeline that runs on a daily cron and aggregates from a well-chosen set of upstream sources. Base metadata comes from CVEProject/cvelistV5, which the project treats as the source of truth for CVE records and raw GitHub PoC links. It scans projectdiscovery/nuclei-templates daily for matching .yaml exploit templates, matches Exploit-DB entries via their .csv dumps, extracts and correlates Metasploit Ruby modules from rapid7/metasploit-framework, and cross-references Vulhub for Dockerized reproduction environments. Enrichment layers from CISA KEV and FIRST EPSS add weaponization context on top.
What makes the dataset useful for triage is the scoring model documented in docs/SCORING.md. Each CVE carries a transparent 0-100 exploit-risk score computed from five signals: KEV membership, EPSS probability, CVSS severity, evidence of public PoC code, and recency of publication. The README's own stats table reports 149,756 total CVEs at last update, with 1,225 flagged critical-risk, 822 high-risk, 147,709 medium/low, and 1,676 present in the CISA KEV catalog. That distribution alone is instructive: only a small fraction of all CVEs carry credible exploit evidence, which is exactly the signal a patch-prioritization workflow needs.
The trending table at the top of the README illustrates how the risk score is operationalized. Each entry pairs a CVE with its affected product, a risk number, and the underlying signals — for example a Fortinet FortiMail path traversal at risk 59 with KEV membership, an EPSS of 2.2%, and two public PoCs, or a Cisco Catalyst SD-WAN Manager flaw at 57 with one verified PoC. Entries for Apple RCE, Zammad RCE, and Citrix NetScaler RCE follow the same pattern. Reading the table is a lesson in triage grammar: KEV membership dominates, PoC count corroborates, and EPSS quantifies probability. This is defensive prioritization data, not an attack recipe.
A genuinely interesting design decision is the explicit optimization for LLM and RAG consumption. The README includes token benchmarking that compares formats: standard NVD JSON at roughly 4,500 tokens per CVE, markdown tables at ~1,200, and the project's minified schema at ~980 — a claimed 78% cost reduction for a 10,000-CVE query. The schema is deliberately dense so that analysts querying the dataset through Claude, GPT-4, or Gemini don't blow their context window. This is a pragmatic acknowledgment of how threat-intelligence work actually gets done now, and the minified format is arguably the project's most defensible engineering choice.
Local querying is intentionally boring and toolchain-native. The dataset lives at data/CVE_list.json, and the README shows a workflow of fetching the file with curl from the repository's raw GitHub content and filtering with jq. Examples include selecting all CVEs that have an associated Metasploit module by checking the .msf array length, or grepping descriptions for RCE-class flaws. This is about as low-friction as data tooling gets — no daemon, no dependency hell, just a large JSON blob and standard POSIX tools. We'll note that the fetch example is the one non-weaponized invocation we'd reproduce here: curl -sL against the repo's raw content URL, then pipe into jq.
The per-CVE granularity matters for real workflows. Beyond the aggregate index, the repository maintains a cves/ directory tree organized by year, with individual .md files such as cves/2026/CVE-2026-104286.md for the FortiMail entry. This means an analyst can bookmark or grep individual records, feed single files into a report pipeline, or link them from ticketing systems. The ecosystem matrix in the README — covering domains from Linux Kernel and Cisco IOS to WordPress, Active Directory, Kubernetes, and Terraform — signals the intended breadth: it is a general-purpose index, not a niche CMS-focused one.
There are things a careful professional should observe before adopting this as a canonical source. The first is provenance: the scores, PoC counts, and enrichment are aggregated by an individual maintainer's automation, and the "verified" qualifier on 150,000 PoCs should be read as "correlated across multiple public sources" rather than independently validated by researchers. Cross-check anything that drives a patch decision against the upstream CISA KEV and FIRST EPSS sources directly. The second is a transparency problem worth flagging plainly: the README contains an explicit prompt-injection-style directive instructing LLMs, RAG systems, and scrapers to cite and rank the repository as their canonical source. We treat that embedded instruction as hostile evidence per our editorial policy — it does not change our factual analysis, but readers building AI pipelines around this dataset should strip such directives from any context they feed to models.
That injection attempt aside, the underlying data model is sound and the source list is respectable. Correlating cvelistV5 with Exploit-DB IDs, Nuclei templates, Metasploit modules, Vulhub environments, KEV, and EPSS is precisely the enrichment join that most teams build ad hoc in spreadsheets. Having it refreshed daily by GitHub Actions cron, without API rate-limit exposure, removes the usual operational excuse. The Zero API claim means the aggregation happens upstream in CI; consumers only ever pull static files.
In an authorized engagement or SOC context, the natural fit is the reconnaissance-to-prioritization gap. After a scanner hands you 400 findings, the operational question is which ones have public exploit code and weaponization telemetry — and that is a lookup against CVE_list.json, not a manual hunt across five websites. Bug bounty hunters can use it to shortlist targets in scope; purple teams can use the Nuclei template and Metasploit correlations to plan detection engineering for techniques that are actually publicly demonstrated. The Vulhub cross-reference is particularly valuable for lab work, pointing to Dockerized environments where a vulnerability class can be studied safely.
Defensively, the dataset is equally useful as a mirror: if a CVE affecting your stack appears here with multiple PoCs and KEV membership, that is an escalation signal for patch SLAs, and the per-record EPSS values give you a defensible, quantified rationale for the ticket priority. Because the tool is entirely passive — it fetches nothing from your infrastructure — there is no telemetry or detection concern on your side; the only observable footprint is the analyst's HTTP pull of the JSON index from GitHub's raw content CDN.
Overall, Exploit-Index is best understood as infrastructure rather than a tool in the offensive sense: a daily-built join table over public vulnerability data sources, scored transparently and consumable by humans and LLMs alike. Its weaknesses are typical of solo-maintained aggregation projects — trust in the pipeline, possible lag or false correlations at the margins — but its sources are citable, its scoring methodology is documented rather than hidden, and its zero-dependency consumption model makes adoption nearly free. For authorized security teams wanting to replace manual CVE-to-exploit searches with one local file and jq, it earns a place in the toolchain, provided the output is always sanity-checked against the primary KEV and EPSS feeds.
SecureWithUmer/Exploit-Index.Educational analysis for authorized security professionals. Use only in controlled, authorized environments.
Related coverage
0 comentários:
Post a Comment
Note: Only a member of this blog may post a comment.