Tuesday, October 6, 2026

T3MP3ST for orchestrating AI agents into an autonomous red team

T3MP3ST for orchestrating AI agents into an autonomous red team

T3MP3ST is a multi-agent offensive-security framework that wraps an AI coding agent you already run into an autonomous recon-to-exploit-to-report pipeline for authorized testing and research.

Toolelder-plinius/T3MP3ST — a multi-agent offensive-security meta-harness that turns existing AI coding agents into autonomous security-testing swarms
CategoryAI-driven red team automation / agent orchestration framework (TypeScript, AGPL-3.0)
Primary UseAuthorized black-box web testing, hint-free CTF solving, white-box source analysis, and coordinated-disclosure vulnerability research, driven from the War Room UI or CLI
Safe UseStrictly for authorized penetration tests, lab environments, CTF sandboxes, and coordinated-disclosure research on systems you own or have explicit written permission to test
Telemetry NoteFully self-hosted and keyless by default; the War Room binds to 127.0.0.1:3333 localhost only, and all mission artifacts (findings, PoCs, disclosure drafts) persist under bench/ paths — defenders observing it in a lab will see heavy coordinated agent traffic against in-scope targets and Docker-staged challenge containers

T3MP3ST, published at elder-plinius/T3MP3ST under AGPL-3.0 and written primarily in TypeScript, is one of the more conceptually interesting entries in the current wave of agent-driven security tooling. Its core bet is stated plainly in the README: the AI coding agent you already run locally — Claude Code, Codex, Hermes, OpenCode, Oh My Pi — is already capable of offensive reasoning, and what it lacks is coordination, an arsenal, and a workflow. T3MP3ST supplies those, positioning itself as the war machine bolted around your existing agent rather than yet another keyed SaaS product. With roughly six thousand stars and a scope table that separates shipped features from scaffolding, it deserves a documentary look for authorized professionals tracking where offensive automation is heading.

The architectural premise is the key differentiator. Instead of provisioning its own model accounts, T3MP3ST treats the locally installed coding agent as the brain and wraps it in a multi-agent harness. There are no new API keys required and no cloud tenant to stand up — the README brands this "keyless warfare." For operators with air-gapped or data-sensitive engagements, the framework also runs fully offline against local models via Ollama, LM Studio, vLLM, or any OpenAI-compatible endpoint configured through TEMPEST_LOCAL_BASE_URL and TEMPEST_LOCAL_MODEL. Notably, tool-calling is driven over text rather than relying on native function-calling, so the Arsenal executes even on models that never shipped that capability.

The kill chain is described as autonomous: recon → exploit → report. Operators interact through a browser-based War Room served at http://127.0.0.1:3333/ui/ or via the CLI, describing a target in plain English to an entity the README calls Op Admiral, which then plans and launches the mission. The README is unusually explicit that the connected agent — not T3MP3ST itself — performs the reasoning, which frames the tool as a meta-harness: orchestration, task decomposition, tool dispatch, and receipt capture layered over an external reasoning engine. A Docker path exists via docker compose up -d, with the container deliberately bound to 127.0.0.1:3333 so the API is not exposed to the network.

What impressed us most, from a credibility-engineering standpoint, is the reproducibility discipline. Every headline number in the README recomputes from committed benchmark data via npm run verify-claims, which the badge claims passes 27/27. The headline benchmark is a 90.1% pass@1 on XBOW's 104-challenge suite — above XBOW's own self-reported 85% — alongside hint-free CTF solves and a "cold hunt" against post-cutoff CVEs the model had never seen during training. This is a meaningful pattern other tool authors should copy: claims that cannot be re-derived from the repository simply do not ship, which sharply distinguishes T3MP3ST from the vibe-driven marketing common in the AI-security space.

The capability matrix is honest about maturity, and professionals should read it carefully before deploying anything against a real scope. Marked stable: black-box web application testing against the XBEN suite, sandbox-jailed CTF solving on Cybench, and a coordinated-disclosure pipeline for robotics, OT, and embedded OSS vulnerability hunting built on OSV records plus live-PoC validation and a refuter stage. White-box source analysis uses what the README calls blind master-builder decomposition, ingesting multiple languages through web-tree-sitter. The riskier domains are explicitly labeled scaffolding: cloud IaC misconfig detection (cloud:bench), mobile static analysis via mobile:bench with an opt-in arsenal including mobsfscan and objection, and binary sink detection (binary:bench) built on decompiler output with ghidra and radare2 available. Smart contracts currently reproduce known Damn Vulnerable DeFi issues rather than discover novel ones.

The opt-in arsenal deserves scrutiny because it reveals the integration philosophy. Rather than reimplementing established tooling, T3MP3ST orchestrates existing binaries: scoutsuite, cloudfox, and pmapper for cloud work, with pacu explicitly gated behind operator action; drozer and gated frida for mobile; gdb, objdump, checksec, and strings for binaries. Gating the more dangerous tools behind explicit consent is a defensible design choice for a framework whose agents can execute commands autonomously, and it suggests the authors have thought about the blast radius of an agent swarm with live tooling attached.

Security posture around the agent boundary is also considered in ways most agent frameworks ignore. Claude Code session reuse is disabled by default because a resumed session may retain provider-side context and ambient tool authority outside what the README calls the Arsenal receipt boundary; operators who trust their configuration can opt in via T3MP3ST_TRUST_CLAUDE_SESSION=1, with the explicit caveat that this is not an Arsenal-enforced sandbox. Timeout knobs (T3MP3ST_LOCAL_AGENT_TIMEOUT_MS, T3MP3ST_TASK_TIMEOUT_MS, T3MP3ST_GENERAL_TIMEOUT_MS) accommodate slow local models. The third-party Novita AI provider is opt-in with a frank disclosure that prompts and metadata leave the machine, and no independent assurance is claimed for sensitive workloads.

Operational hygiene extends to the update mechanism. npm run update performs an interactive sync from upstream with y/N confirmation, backing up protected paths listed in scripts/update-protected.txt before touching the tree; npm run update:dry is a read-only preview safe even on tarball installs, and npm run update:hard is an opt-in hard reset that still restores protected files. The protected list is itself informative: .env files, .keys.local, bug-bounty platform credentials in .keys.bounty.json, expensive benchmark corpora under bench/cybench/, and — most interesting for researchers — bench/wild-hunt/, where cold-hunt findings, PoCs, and disclosure drafts accumulate as pre-coordination vulnerability material. That last path tells you the intended output of a real campaign is a responsible-disclosure package, not a loot bag.

For defenders, T3MP3ST is worth studying as a preview of adversary automation. A coordinated swarm doing recon, exploitation, and reporting against a target will produce recognizable patterns: high-velocity structured reconnaissance, consistent agent-authored request signatures, and staged containerized challenge infrastructure in lab settings. The fact that the framework self-hosts, defaults to localhost binding, and leaves rich local audit artifacts in bench/ means blue teams replicating it in a range get full telemetry of what an autonomous offensive campaign looks like from the inside — valuable for building detection content against the behavior class rather than specific tool signatures.

Installation is conventional for a Node.js project: npm install followed by npm run server brings up the War Room locally, and offline operation adds an ollama serve and two environment variables. This article documents the tool for educational purposes; the README itself opens its legal section with a blunt authorized-use-only warning, noting that unauthorized access is illegal in most jurisdictions and that the operator alone bears responsibility for staying inside rules of engagement. Anyone evaluating T3MP3ST should point it exclusively at owned systems, lab sandboxes, CTF infrastructure, or programs with explicit written permission.

The bigger picture is that T3MP3ST is a credible artifact of a genuine shift: offensive capability migrating from bespoke tooling into orchestration layers over general-purpose coding agents. Its unusual virtues — re-derivable benchmarks, an honest status table, gated dangerous tools, and disclosure-oriented output — make it a better-behaved exemplar of the genre than most. Watch the scaffolding rows (cloud, mobile, binary) as they mature toward benchmarked status; if the cold-hunt pipeline on post-cutoff CVEs keeps producing coordinated-disclosure findings, this repository will remain one of the more consequential things to monitor in authorized offensive research.

Official project repository for elder-plinius/T3MP3ST.
Download Tool

Educational analysis for authorized security professionals. Use only in controlled, authorized environments.

Share articleFacebookXLinkedIn

Continue exploring

Browse all articles →

0 comentários:

Post a Comment

Note: Only a member of this blog may post a comment.