Wednesday, September 30, 2026

TRACE for probing security risks in LLM agents

TRACE for probing security risks in LLM agents

TRACE is an academic red-teaming framework that decomposes complex safety test cases into benign-looking subtasks and uses an evolving strategy library to evaluate LLM agents in authorized research settings.

ToolZJU-LLM-Safety/TRACE — task-aware adaptive self-evolving framework for agentic jailbreak research
CategoryAI/LLM safety evaluation and red-team research framework
Primary UseStress-testing LLM-driven coding agents such as Claude Code and Codex against decomposed multi-step safety cases in controlled benchmarks like AdvCUA
Safe UseIntended exclusively for authorized AI-safety researchers and defenders evaluating their own agents in isolated lab environments; not for attacking third-party systems
Telemetry NoteRuns leave JSONL result files, strategy-library state, and agent session trajectories; on the defensive side, its characteristic pattern — long multi-turn interactions building plausible execution contexts — is observable in agent transcripts and refusal logs

TRACE, published by the ZJU-LLM-Safety group as ZJU-LLM-Safety/TRACE, is a research framework aimed at one of the harder problems in modern AI safety: evaluating agentic systems that combine an LLM planner with tool execution. Unlike classic single-prompt jailbreak studies, TRACE models the attack surface the way an actual misuse scenario would unfold — as a sequence of individually innocuous steps that only become harmful in aggregate. The repository is written in Python and, at the time of this writing, carries roughly 30 stars, which is consistent with a fresh academic release accompanying ongoing research rather than a production tool.

The core architectural idea is scheme-based task decomposition. The improved_decomposer/ module holds around twenty procedural decomposition schemes that transform a complex harmful test task into sequences of comparatively benign, simple subtasks. Crucially, the framework does not pick a decomposition arbitrarily: candidate sequences are scored along two measured dimensions — harmfulness and difficulty — and the optimal sequence is selected for instantiation. This scoring loop is what distinguishes TRACE from naive prompt-splitting approaches, because it treats decomposition itself as a search problem with an explicit objective function.

The second pillar is execution-oriented multi-turn interaction. When a subtask triggers an execution refusal from the target agent, TRACE builds a structured subtask profile and retrieves a matching interaction strategy from a strategy library stored under experiment/strategy_library/. The selected strategy is used to instantiate the subtask and gradually assemble a plausible execution context over multiple turns, coaxing the agent toward completing the step. For defenders, this is the most instructive mechanism in the codebase: it demonstrates concretely how refusal events become feedback signals that shape subsequent conversational framing.

The third pillar, adaptive recovery and strategy evolution, is where the framework earns its "self-evolving" label. When an intermediate step fails, TRACE localizes the failing subtask, preserves prior progress, adapts the strategy, and resumes from the point of failure rather than restarting the whole chain. Execution feedback is continuously folded back into the strategy library, meaning the framework's effectiveness compounds across runs. The split between experiment/strategy_library/ for storage, retrieval, and revision, and experiment/main.py for orchestration and feedback, makes this loop explicit in the repository layout.

The repository is cleanly organized for a research artifact. experiment/runtime/ hosts the runtime components for models, strategies, and subtask sequences; experiment/config/ holds configurations and prompt templates; agent_execution/ contains the target-agent interfaces, session handling, trajectory collection, and verification logic; and tools/ provides benchmark bridges, API adapters, and diagnostics. Two recorded demonstrations targeting Claude Code and Codex are shipped in assets/ as claude_demo.mp4 and codex_demo.mp4, which gives reviewers a way to understand the interaction dynamics without running anything themselves.

Environment setup is straightforward for an authorized lab. The specification targets Linux and Python 3.11, declared in environment.yml, and a conda environment is the intended install path. The two relevant commands are the environment creation and the decomposition entry point, which is where any evaluation begins.

Setup reduces to conda env create -f environment.yml followed by conda activate trace. API access is configured through standard OPENAI_API_KEY and OPENAI_BASE_URL environment variables pointing at any OpenAI-compatible Chat Completions endpoint, which keeps the framework model-agnostic on the attacker side. The decomposition step is then invoked via python improved_decomposer/improved_task_decomposer.py with a tasks file, an output path, a model name, the number of candidate decompositions per task, and the --use_schema_decompose flag to enable scheme-based candidate generation; --help on both entry points documents the full option surface.

On the experiment side, the main.py orchestrator accepts a benchmark dataset with precomputed subtasks, a config.yaml, a lab backend, a Docker Compose file for the isolated target environment, an attacker model path, a target backend selector such as codex, and a strategy-library file — output lands in JSONL result files under outputs/. The presence of a dedicated --compose-file argument and benchmark harnessing in tools/ signals that the authors intend everything to run inside a self-contained lab, with the AdvCUA benchmark as the reference evaluation. That isolation-first design is the correct posture for research of this kind, and operators reviewing the code should preserve it.

From a defensive standpoint, TRACE is more valuable as a description of adversary behavior than as anything to deploy. Blue teams building guardrails for coding agents can mine the strategy-library design for detection heuristics: multi-turn sessions that incrementally construct execution contexts, subtask sequences whose individual harm scores are low but whose composition is not, and resumed sessions that probe a previously refused step under a reframed strategy. Agent-side logging that captures full trajectories — exactly what agent_execution/ collects on the research side — is the telemetry defenders need on theirs.

There are caveats worth stating plainly. The framework is an offensive evaluation instrument by design; pointing it at agents you do not own or operate is misuse, and the academic framing in the README assumes institutional review and controlled targets. The absence of a declared license in the repository metadata means potential adopters should confirm redistribution terms with the authors before integrating any of the decomposition or strategy code into their own evaluation pipelines. There is also no CI or test suite visible in the structure table, so expect research-grade code quality.

Within its niche, TRACE sits alongside a growing body of agentic-safety evaluation work, and its distinguishing contribution is the closed feedback loop: decomposition candidates scored on harmfulness and difficulty, refusal-driven strategy retrieval, checkpointed recovery, and library evolution. For teams that ship LLM agents with tool access, studying how this loop operates — and then testing whether their own refusal and monitoring layers withstand gradual context construction — is a legitimate and increasingly necessary part of the pre-release security review. The repository serves that purpose well, provided it stays in the lab where it belongs.

Official project repository for ZJU-LLM-Safety/TRACE.
Download Tool

Educational analysis for authorized security professionals. Use only in controlled, authorized environments.

Share articleFacebookXLinkedIn

Continue exploring

Browse all articles →

0 comentários:

Post a Comment

Note: Only a member of this blog may post a comment.