
pyvex wraps Valgrind's VEX IR in Python bindings, letting authorized researchers translate machine code from many architectures into a uniform representation for static binary analysis.
| Tool | angr/pyvex — Python bindings for Valgrind's VEX intermediate representation |
| Category | Binary analysis / IR lifting library |
| Primary Use | Uplifting raw machine code into IRSB blocks via pyvex.lift() for architecture-independent static analysis, typically as part of angr workflows |
| Safe Use | Intended for authorized security research, vulnerability research on firmware and binaries you own or are licensed to assess, and academic program analysis |
| Telemetry Note | Purely local analysis library; it performs no network activity and leaves no artifacts beyond Python process state — defenders would never observe its use on a target |
pyvex is one of the load-bearing components of the angr binary analysis ecosystem: Python bindings around the VEX intermediate representation originally built for Valgrind. Where most reverse-engineering tooling forces you to reason about x86, ARM, MIPS, or PPC semantics separately, pyvex lets you lift a basic block of machine code into a single, architecture-agnostic IR and then write one analysis that works everywhere. The repository, angr/pyvex, is Python under a BSD-2-Clause license, published on PyPI as pyvex, and documented at api.angr.io alongside the rest of the angr project suite.
The core API is compact. You call pyvex.lift() with raw bytes, a base address, and an archinfo.ArchAMD64()-style architecture descriptor, and you get back an IRSB — an IR Super Block. The README's example lifts five 0x90 NOP bytes at 0x400400 on AMD64. From that point the block is inspectable Python: irsb.pp() pretty-prints it, irsb.next exposes the IR expression for the unconditional jump target, and irsb.jumpkind tells you the exit type — a call, a return, a syscall, or a plain boring jump. Iteration over irsb.statements gives you statement-level access for custom passes.
Installation is the standard one-liner, pip install pyvex, which pulls prebuilt wheels bundling the compiled VEX translation core — there is no need to build Valgrind yourself. That matters because the heavy lifting, decoding arbitrary instruction sets into IR, happens in C, and the bindings expose the result as native Python objects with .pp() methods on nearly everything for quick debugging during interactive analysis sessions.
To understand why pyvex exists, it helps to look at what the README says an IR buys you. Four classes of architecture difference are abstracted away: register naming, memory access (including ARM's dual endianness modes expressed as LDle and LDbe), x86-style memory segmentation through segment registers, and instruction side effects such as Thumb-mode flag updates or stack pointer movement on push and pop. The last point is the quiet killer feature — side effects are made explicit in the IR rather than being implicit folklore you have to remember per-architecture, which is exactly the kind of detail that silently breaks hand-rolled disassembly-based analyses.
The VEX object model has five tiers. IR Expressions represent values — constants, register reads like GET:I32(16), memory loads, arithmetic operations like Add32, if-then-else constructs, and calls to C helper functions used for things like condition-flag computation. IR Operations modify expressions. Temporary variables, numbered from t0 and strongly typed as 64-bit integer or 32-bit float, serve as the IR's internal registers. IR Statements model state changes — WrTmp, Put for register writes, STle/STbe for stores, and Exit for conditional branches out of the middle of a block. Finally, an IRSB bundles the statements of one extended basic block, which may have several exits.
The README's worked ARM example makes the mechanics concrete. The single instruction subs R2, R2, #8 becomes five IR statements: the register is read into t0, the constant 0x8:I32 is materialized, Sub32 computes the difference into t3, the result is written back with PUT(16), and the final statement increments the program counter to 0x59FC8. Notice that even the PC update is explicit — nothing in VEX is implicit state mutation, which is what makes symbolic execution engines like angr feasible on top of it.
The register model is worth internalizing early because it surprises newcomers: registers are modeled as a separate memory space addressed by integer offsets, so AMD64's rax lives at offset 16 in that space rather than being referenced by name. The README points to libvex_ir.h in the angr/vex repository as the authoritative documentation, and irsb.tyenv.types exposes the type environment mapping each temporary to its type, indexed the same way — irsb.tyenv.types[0] gives you the type of t0. If you're building analyses that need to reason about data widths, this is where you look.
An important caveat stated plainly in the README: this is a syntactic representation only. The IRSB tells you what a block means — which registers are read, what gets stored where — but carries no runtime context, so you cannot say what actual data a store instruction writes without supplying a state model yourself. That separation of concerns is deliberate; pyvex is the translation layer, and semantic reasoning (concrete or symbolic) is the job of higher layers in the angr stack such as angr.Project and its simulation engine.
In an authorized workflow, pyvex shows up wherever you need custom lifting or IR-level inspection: writing targeted taint-style checks over IRStmt.Store statements, extracting call graphs by scanning jumpkind values, normalizing firmware blobs from mixed architectures into one comparable form, or building concolic tooling on top of angr. The store-walking and exit-inspection snippets in the README — iterating irsb.statements, filtering isinstance(stmt, pyvex.IRStmt.Store) or IRStmt.Exit, and pretty-printing stmt.data, stmt.guard, and stmt.dst — are the idiomatic starting points for that class of code.
The project's provenance is academic and legitimate: the citation block references the Firmalice paper on automatic detection of authentication bypass vulnerabilities in binary firmware (NDSS 2015), which is the research lineage that produced the whole angr platform. This is analysis infrastructure, not offensive tooling — it touches nothing over the network, runs entirely in-process, and is equally useful to defenders building malware triage pipelines and to compiler engineers studying code generation.
For defenders, that dual-use framing matters less than usual here, because pyvex is essentially a library for reading code, not acting on systems. Its value in a blue-team context is normalization: turn a stripped multi-architecture malware sample into IR statements, diff the statement streams against known families, and detect subtle code-reuse variants that byte signatures miss. At roughly 380 stars it is a mature, purpose-built dependency rather than a flashy standalone tool, and that steadiness — PyPI releases, BSD licensing, deep documentation in libvex_ir.h — is exactly what you want in a foundation layer for serious binary analysis work.
angr/pyvex.Educational analysis for authorized security professionals. Use only in controlled, authorized environments.
Related coverage
0 comentários:
Post a Comment
Note: Only a member of this blog may post a comment.