Sunday, October 4, 2026

Automating full-pipeline API attack surface mapping with APIHarvester

Automating full-pipeline API attack surface mapping with APIHarvester

APIHarvester is a Python black-box API security scanner that chains subdomain discovery, endpoint enumeration, and OWASP API Top 10 test simulations into a single authorized assessment run.

Toolpiratesshield/APIHarvester — full-pipeline black-box API security scanner written in Python
CategoryAPI attack surface discovery and authorized security testing automation
Primary UseAutomating endpoint enumeration, parameter discovery, and OWASP API Top 10 (BOLA, BFLA, SSRF, mass assignment) checks during penetration tests of APIs you are contracted to assess
Safe UseEducational and documentary analysis for authorized professionals; run only against APIs in your own lab, staging environments, or explicit-scope penetration test engagements
Telemetry NoteGenerates high request volume (--burst of 20+ rapid requests for rate-limit checks, 20 threads by default), leaving dense access-log signatures and structured output/{target}_{timestamp}/ artifact directories; WAF detections are logged per host in waf_results.jsonl

APIHarvester, hosted at piratesshield/APIHarvester, is a Python-based black-box scanner that tries to compress an entire API assessment workflow into a single invocation. The README describes a full pipeline: starting from a bare root domain, it enumerates subdomains, discovers live HTTP hosts, crawls for endpoints including dynamic SPA/XHR routes, identifies parameters, probes allowed HTTP methods, and then runs a battery of attack simulations mapped to the OWASP API Top 10. With 32 stars and a stdlib-first design philosophy, it positions itself as a grab-and-run tool for practitioners who want coverage without a heavy dependency chain.

The most interesting architectural decision is the standard-library-only core. The README states that all core scripts are designed to be stdlib-only, and that apiharvester automatically falls back to pure-Python implementations when external tools or libraries are absent. This guarantees the scanner executes out of the box on a bare Python 3 installation, which matters in restricted assessment environments where installing packages is awkward. Optional accelerators exist — scripts/install_requirements.sh fetches SecLists wordlists and Kiterunner route schemas and can install helper binaries via go install and pip — but they are enhancements rather than prerequisites, and scripts/check_requirements.sh performs a read-only audit of what is present.

Repository layout tells you how the project is organized. The main package lives in apiharvester/ and runs as python3 -m apiharvester, while apiharvester.py and apisec.py are standalone single-file distributions of the scanner — useful for dropping onto a jump host. Two support scripts stand out: api_deep_discovery.py, a dynamic crawler that borrows Katana headless-browser code for discovering endpoints that only materialize when JavaScript executes, and api_intelligence_engine.py, described as the pipeline aggregator and passive vulnerability classifier. The payloads/ directory ships reconnaissance data: params.txt with 25,889 parameter-name candidates, directories.txt with 62,281 API path patterns, subdomains.txt with 5,000 subdomain variants, and kiterunner/ route schema files for accelerated endpoint enumeration.

The dual-token design is where the tool gets operationally clever for authorization testing. The --auth flag carries a high-privilege bearer token while --auth2 carries a low-privilege one, and the differential between the two sessions is exactly what modules like bola and bfla need to prove cross-account access or privilege escalation. Without that pair, several attack phases simply degrade or skip, which is a sensible fail-quiet behavior. Thread pool size defaults to 20 via --threads, HTTP timeout to 10 seconds, and the --burst count for rate-limit verification defaults to 20 rapid requests — operators should tune these against the capacity of the target environment they are authorized to test.

Attack coverage maps directly onto the OWASP API Security Top 10 (2023). bola handles Broken Object-Level Authorization using an improved ID-candidate generator, _generate_id_candidates(), which produces 5-13 variants per endpoint: numerics like 0, 1, 99, ID±1, UUID last-segment flips such as 00000000 and ffffffff, common hex/Mongo patterns, and strings like admin and guest. broken_auth covers unauthenticated endpoint discovery, JWT weak-secret cracking, alg=none bypass attempts, claim tampering, kid injection, and OPTIONS/HEAD method bypasses — a class of misconfiguration the README explicitly notes arises when servers apply authentication only to GET/POST. mass_assignment injects privilege-escalation fields such as role, is_admin, and verified into PUT/PATCH bodies, and bfla probes sensitive paths like /admin and /impersonate with and without the low-privilege token.

Beyond the Top 10 mapping, three bonus modules extend the surface. injection covers error-based and time-based blind SQL injection, XSS, and command injection. secrets pattern-matches response bodies for leaked credentials — AWS AKIA keys, Google AIza keys, xox Slack tokens, Stripe live keys, GitHub tokens, private key blocks, JWTs, and generic api_key= / password= assignments. reliability is described as RESTler-style fuzzing: boundary query values (huge numbers, null bytes, oversized strings), malformed JSON bodies (wrong top-level types, deep nesting, huge arrays), flagging any 5xx response as a reliability bug independent of the OWASP taxonomy. That reliability angle is a genuinely different lens — crash discovery rather than categorization — and echoes the academic RESTler lineage the README cites as inspiration.

Output handling is one of the stronger engineering choices here. Every scan writes recon artifacts to a timestamped directory, output/{target}_{YYYYMMDD_HHMMSS}/, before any attack phase runs. Files include fqdn.txt for discovered subdomains, fqdn_active.txt for live hosts, fqdnwithendpoint.txt for endpoint URLs, withparam.txt for parameterized URLs, objectshape.txt mapping response JSON field names per endpoint, waf_results.jsonl with per-host WAF and JS-challenge detections, endpoint_methods.jsonl for allowed HTTP verbs, and swagger_specs/*.json for any discovered OpenAPI documents. Everything is line-delimited plain text or JSONL, which means the directory doubles as a portable recon dataset that pipes cleanly into grep, jq, httpx, nuclei, or ffuf without a custom parser — the README even shows recipes for seeding ffuf wordlists from discovered parameters.

The phase-separation flags make iteration cheap. --skip-recon reuses existing output files, --recon-dir loads a prior run's recon data, and --attacks-only runs just the attack phases against it — useful when you have already mapped the surface and want to re-test after a fix cycle. --attacks accepts a comma-separated list (bola,secrets, for example) from the full set of bola, broken_auth, mass_assignment, rate_limit, bfla, business_logic, ssrf, misconfiguration, inventory, sspp, injection, reliability, secrets. Reporting comes in two forms: an interactive HTML dashboard via --html and line-delimited JSON findings via --json, both of which suit handoff to a client report.

Installation is straightforward and non-weaponized: clone the repository, optionally run ./scripts/check_requirements.sh to audit prerequisites and ./scripts/install_requirements.sh to fetch the SecLists-derived wordlists, then pip3 install -r requirements.txt for the optional Python accelerators. Invocation is a single module call such as python3 -m apiharvester example.com --html report.html, with authentication tokens supplied only for targets you are explicitly authorized to test. Note that the tool uses a permissive TLS context by default so self-signed staging certificates do not break scanning — convenient in labs, but a behavior worth knowing about before pointing it anywhere sensitive.

From a defensive standpoint, this tool is equally valuable as a mirror: blue teams can study its module list as a checklist of what automated API attackers will try. The rate-limit module flags endpoints that return 200 instead of 429 under 20+ rapid requests, the misconfiguration module sends untrusted Origin headers to test CORS, and the secrets module hunts for credential patterns in responses — all detectable behaviors. The dense, patterned request traffic (20 threads, 62k-path wordlists, parameter fuzzing) leaves clear signatures in API gateway logs and will trip most WAFs, which is precisely why the scanner records its own waf_results.jsonl detections. Defenders running APIHarvester against their own staging systems get a prioritized list of exactly those gaps before an adversary's automation finds them.

Caveats worth weighing before adopting it: the repository has no license declared, which complicates commercial use; there are no topics or stated versioning; and the README's improvement section suggests the project iterated quickly from an earlier, less capable version. The breadth of automated attack simulation also demands strict scope discipline — mass assignment injection, JWT tampering, and SSRF probing are exactly the categories of activity that must never touch systems outside a signed engagement. Used inside that boundary, APIHarvester is a legitimately comprehensive attempt to collapse recon, discovery, and OWASP API Top 10 validation into one repeatable pipeline, and its plain-text, tool-agnostic output format makes it a good citizen in a larger assessment toolchain.

Official project repository for piratesshield/APIHarvester.
Download Tool

Educational analysis for authorized security professionals. Use only in controlled, authorized environments.

Share articleFacebookXLinkedIn

Continue exploring

Browse all articles →

0 comentários:

Post a Comment

Note: Only a member of this blog may post a comment.