
REx@Skill orchestrates seven Claude subagents and a pinned reverse-engineering toolchain to find, prove, and rule out defects in compiled binaries during authorized assessments.
| Tool | tihanyin/REx-skill — an agentic reverse-engineering execution skill for Claude Code that coordinates seven subagents and ~20 pinned tools to find and prove binary defects |
| Category | Agentic binary analysis / vulnerability research orchestration |
| Primary Use | Running /re-analyze on a binary or corpus during authorized engagements to produce findings, ruled-out hypotheses, and environment limitations with named sources, sinks, and broken guards |
| Safe Use | Authorized penetration tests, code audits of software you own or are contracted to assess, CTF practice, and defensive vulnerability research in isolated lab environments |
| Telemetry Note | Purely local: install.sh writes only to ~/.claude/, the Nix devshell only mutates PATH, and analysis runs on the host — nothing phones home; defenders observing it would see the pinned toolchain processes (Ghidra, radare2, AFL++, qemu-user) executing in the analysis sandbox |
REx@Skill is a Claude Code skill — version 1.0.0, MIT-licensed, written around Python 3.12/3.14 — that turns agentic LLM workflows loose on compiled binaries with a methodology borrowed from serious vulnerability research: nothing counts as a finding unless it is proven. The project sits at roughly 59 stars and markets itself with a deceptively simple promise, "find the defects in a compiled binary — and prove them," backed by a specific architectural claim: seven subagents, one shared evidence tree, and a fully pinned toolchain so that two runs are actually comparable. That last clause is the tell that this is an operator's project rather than an LLM demo — reproducibility of decompiler output is a real pain point in collaborative audits.
Installation is deliberately minimal. The skill ships as one installer that places a single skill, 12 references, 7 subagents, and 33 pipeline scripts under ~/.claude/ — it never installs Claude Code for you, backs up anything it would overwrite, and supports install.sh --uninstall for clean removal. You can also git clone https://github.com/tihanyin/REx-skill and run ./install.sh --prefix to target an alternate location. Once installed, the entry point is a claude session with /re-analyze path/to/binary for a single target or /re-analyze path/to/directory/ to process a whole corpus with evidence gathered in parallel.
The tool inventory reads like a checklist of the state of the art in binary analysis. Ghidra 12.1.2 does primary decompilation, with radare2 6.2.0 and rizin 0.9.1 as independent second opinions to cross-check the first decompiler. qemu-user 11.1.0 with nine cross-sysroots executes foreign-architecture binaries, angr 9.2.154 reasons about input reachability, z3 4.16.0 discharges arithmetic claims like out-of-bounds indices or zero divisors, and AFL++ 5.00c generates crashing inputs at scale. Around these sit valgrind, libdislocator, capa, floss, Unicorn, Triton, Frida, semgrep, cppcheck, and pwntools, each pinned to an exact version the project was built and measured against.
What elevates the design is the honest handling of absence. Every tool is optional: scripts/capabilities.sh reports what the host actually has, each of the 33 scripts names the tool it cannot find, and a missing capability narrows the output into an explicit limitations section rather than failing silently or hallucinating coverage. In an agentic context this matters more than in traditional pipelines — an LLM agent that quietly loses its solver can still produce confident prose about arithmetic safety, and REx@Skill's insistence on declaring what the host could not run is a direct mitigation of that failure mode.
The subagent architecture is the intellectual core of the project. Agent 01 (re-recon) runs first and alone, converting the binary into shared evidence — decompiled C, disassembly, strings, a reachability map, and runtime behavior — and explicitly does not hunt bugs, because a confident finding from the recon agent is identified as a failure mode. Agents 02 through 06 then read that evidence concurrently, each with a distinct mandate: re-bughunt builds the strongest honest case that defects exist and enumerates every sink; re-safety tries to prove the binary sound and reports every obligation it cannot discharge; re-arithmetic handles size, index, width, and signedness across call boundaries, discharged by z3 rather than in someone's head; re-lifecycle covers allocation, ownership, initialization, and the error paths nobody tested; re-logic examines authentication and state machines.
The most interesting design decision is isolation between the readers. The five analysis agents cannot see each other's findings while working, because — as the README bluntly puts it — five readers sharing one decompiler make the same mistakes. Keeping them apart buys five independent readings instead of one opinion repeated five times. A seventh agent then reads the code first and the other agents' findings second, acting as an adjudicator over which claims hold up. This is a genuinely thoughtful answer to the biggest weakness of LLM-driven analysis: correlated errors masquerading as consensus.
The evidentiary standard is where REx@Skill distinguishes itself from generic "AI finds bugs" tooling. Nothing is reported as a bug unless four elements are named and located in the binary: where attacker-controlled data enters (the source), the operation it can break (the sink), the check that should have stopped it and didn't (the broken guard), and who is harmed (the affected principal). Miss any one of those and the item ships as an unproven lead, not a finding. Every run also delivers what was ruled out and why — negative results as first-class output — plus the environment's stated limitations. For an authorized assessment report, that ruled-out section is often as valuable as the findings themselves.
Reproducibility is treated as an architectural requirement, not an afterthought. The devshell uses a Nix flake to install the entire toolchain pinned to exact versions — tested on Ubuntu 22.04.5 LTS with Determinate Nix 3.22.4 — and three locked revisions (nixpkgs, nixpkgs-angr, and capa-rules v9.4.0, with devshell/flake.lock as the authority) rebuild the whole environment on any machine at any future point. The README's rationale is worth internalizing: the decompiler's output is the analyst's input, so two analysts on two Ghidra versions are not doing the same experiment, and neither can check the other's result. Entering the shell changes PATH and nothing else; exit restores the system exactly.
Architecture coverage is broad and explicitly uncapped. The skillset has been tested across ten architectures — x86-64, i686, ARM, AArch64, MIPS and MIPS64 in both endiannesses, 32-bit PowerPC, RISC-V, and Apple arm64 — across ELF, PE, and Mach-O formats, firmware images, and raw blobs. The README is careful to state that nothing caps the list there: any architecture your decompiler can lift is in scope, which is the correct boundary condition for a tool that depends on Ghidra headless analysis as its evidence generator.
For authorized workflows, this tool slots in as a force multiplier at the triage stage of a binary audit — assessing your own builds, contracted third-party code review, firmware security research, or CTF preparation. It is not an exploitation framework; pwntools and Frida appear in the toolchain for offset calculation, ELF parsing, and runtime observation during analysis, but the product of a run is evidence and reasoned findings, not weaponized payloads. The four-part source/sink/guard/principal requirement keeps output anchored to what is provably in the binary rather than speculative severity inflation.
Caveats for prospective users: the project depends on Claude Code as its execution substrate, so LLM licensing and API costs apply, and the quality of decompiled-C evidence gates everything downstream — the skill's own framing acknowledges this by pinning decompiler versions obsessively. The one-line installer fetches and executes a remote install.sh, so review it first as you would any curl-pipe-to-shell pattern, or use the git clone path and read install.sh before running it. Within those bounds, REx@Skill is one of the more methodologically serious entries in the young agentic-security space: less a magic bug finder than a disciplined orchestration layer that makes a proven toolchain and a falsification-oriented review process available to authorized analysts.
tihanyin/REx-skill.Educational analysis for authorized security professionals. Use only in controlled, authorized environments.
0 comentários:
Post a Comment
Note: Only a member of this blog may post a comment.