
security-harness orchestrates a multi-agent static review pipeline — mapping, hunting, chaining, verifying, and reporting — for security professionals working on code they own or are authorized to test.
| Tool | dmdhrumilmistry/security-harness — a multi-agent application-security review harness built as a plugin for Claude Code, Gemini CLI, opencode, and Codex |
| Category | Static application security analysis via LLM agent orchestration (Python, MIT-licensed) |
| Primary Use | Authorized static source review: mapping a codebase, hunting vulnerability classes with per-class knowledge bases, chaining findings, verifying impact, and emitting SARIF/JSON/PDF reports |
| Safe Use | Strictly for codebases you own or are explicitly authorized to test — the README states the harness performs static analysis only and does not attack live third-party systems |
| Telemetry Note | All artifacts land under <target>/.security-harness/<run-id>/; the PR-review skill writes inline comments and a security/pr-review commit status but opens no issues or external trackers; a local sh-review-cache stores run state in the OS cache directory, never in the repo |
security-harness is best understood as a plugin marketplace rather than a single tool: a repository whose plugins/security-harness/ tree packages skills, agents, and shared contracts into an installable unit for Claude Code, with adapters for Gemini CLI, opencode, and Codex. The architecture is deliberately hub-and-spoke. A single router skill, sh-router, accepts natural-language appsec requests and dispatches them to sh-security-review, the pipeline orchestrator that runs Stages 0 through 5. Behind it sit fifteen sh-kb-* knowledge bases — one per vulnerability class — and five specialized agents: sh-recon, sh-hunter, sh-chainer, sh-verifier, and sh-reporter.
The pipeline reads like a manual senior reviewer's workflow encoded as state machine. Recon builds a map: a Graft code graph plus stack detection, SBOM generation, and CVE enumeration. Hunt runs source-to-sink tracing, one hunter per vulnerability class, in parallel. Chain combines discrete findings into plausible escalation paths. Verify — described in the README as the precision gate — applies offensive, security-engineering, and developer lenses before confirming real impact. Report emits README.md, findings.json, results.sarif, report.html, report.pdf, and report.docx. Working state persists as findings.jsonl, chains.md, and verified.jsonl under <target>/.security-harness/<run-id>/.
The vulnerability-class coverage is broad and explicitly enumerated: access-control (IDOR/BOLA/privilege escalation), sqli, xss, ssrf, command/code/SSTI/LDAP injection, auth (session/JWT), deserialization, path-traversal, secrets, csrf, xxe, open-redirect, crypto, race-conditions, and file-upload. Dependency CVEs and SBOM data are folded into the recon stage rather than treated as a separate hunt class, which is a sensible structural decision — dependency issues are enumerated from manifests, not discovered by data-flow tracing.
Installation is refreshingly low-friction for Claude Code: /plugin marketplace add dmdhrumilmistry/security-harness followed by /plugin install security-harness. For Gemini CLI, gemini extensions install https://github.com/dmdhrumilmistry/security-harness uses a manifest at the repo root. For any agentskills.io-compatible agent, python3 scripts/sync-agent-skills.py --install agents copies the skills into ~/.agents/skills. The marketplace command also accepts full git URLs and local clone paths, which matters for air-gapped or enterprise-review contexts.
The most interesting internal decision is dependency bootstrapping. Stage 0 auto-installs what's missing — @nanonets/graft via npm install -g, then SBOM tooling (syft) and CVE scanners (preferring grype, falling back to trivy or osv-scanner), plus wkhtmltopdf or pandoc for document output. Crucially, every install is announced, prefers no-elevation methods, and never blocks the run: missing capabilities are marked unavailable and the pipeline degrades gracefully to native search and manifest parsing. That fail-soft posture is what separates a harness designed for real heterogeneous environments from one that shatters on the first missing binary.
Cost control is equally deliberate, and it is the default rather than an option. Each stage runs on a model matched to its cognitive load: haiku for mechanical pattern classes like secrets and csrf, sonnet for source-to-sink tracing classes like sqli and ssrf, and opus for deep-logic classes like access-control and race-conditions. The verify and chain stages stay on opus because precision lives there. Overrides are granular — models:verify=opus,hunt=sonnet targets individual stages, models:max flattens everything to opus, models:cheap downshifts aggressively at some precision cost. Further savings come from spawning hunters only for classes with real attack surface and having hunters query the Graft graph instead of reading whole files; findings.json and SARIF are generated by a deterministic script, not burned model tokens.
Invocation is natural-language through the router — /sh-router full security review of ./api or /sh-router find SQLi and IDOR in src/ — or direct with parameters: /sh-security-review . classes:sqli,access-control,ssrf depth:deep, and stage:report regenerates deliverables from the latest run without re-hunting. Each published finding carries a payload, a PoC, the verification verdict, CWE/OWASP identifiers, a CVSS score, and a code-level mitigation — the schema living in references/ alongside the SARIF mapping and rubrics. Note the scope boundary stated in the README itself: static analysis on code you own or are authorized to test, no attacks against live third-party systems.
The sh-pr-review skill is where the project gets genuinely opinionated about DevSecOps ergonomics. It reviews a single pull request rather than a whole codebase, clones the target to a temporary directory because hunters read full files rather than just the patch, and posts inline comments plus a security/pr-review commit status that branch protection can enforce. Phase 7 asks before posting anything, and it pre-checks write access to the target repo so an unauthorized run fails in seconds rather than after ten minutes of analysis. The review event is always COMMENT, never REQUEST_CHANGES or APPROVE — the commit status is the sole enforcement surface.
Three design rules make it usable as a merge gate. Only findings with pr_impact of introduced or aggravated block a merge; pre_existing issues are reported but never fail the check, and when a hunter cannot distinguish aggravated from pre_existing it must conservatively pick the latter. Depth follows risk: an orchestrator-level triage maps changed paths and added-line sink tokens onto the same class slugs the knowledge bases use, and Tier 0 — a PR with no security-relevant change — launches zero subagents yet still sets the status. Re-pushes don't spam because each comment carries a fingerprint computed without line numbers, so re-reviews add only what is new and list resolved items separately.
Incrementalism is engineered at three layers. Deduplication happens before spending, not before posting: fingerprints already on the PR are read in Phase 1 and handed to hunters and the verifier, so the most expensive model never re-confirms a conclusion already written on the PR. A local sh-review-cache stores file hashes, findings, and verdicts in the OS cache directory — never in the repo, since cross-repo reviews run in deleted temp clones — enabling re-hunting of only changed files. And a base-branch merge that leaves PR files byte-identical re-stamps the previous verdict onto the new SHA with no agents launched, guarded by a conservative check that the base delta touches nothing the findings depend on. The cache key hashes every sh-kb-* knowledge base, so a KB update invalidates prior conclusions — a stale entry in a security tool, as the README puts it, is not slow but wrong.
For defenders, the telemetry picture is clean and predictable: everything local lands under .security-harness/, the graft/ graph directory is auto-gitignored in the target, and the only external writes are PR comments and a commit status. For authorized assessment teams and AppSec programs, security-harness represents a credible pattern for AI-assisted review: tiered model economics, verified-findings-only reporting, SARIF integration into existing developer workflows, and enforcement semantics designed to survive contact with real engineering teams rather than get disabled in week two.
dmdhrumilmistry/security-harness.Educational analysis for authorized security professionals. Use only in controlled, authorized environments.
Related coverage
0 comentários:
Post a Comment
Note: Only a member of this blog may post a comment.