Sunday, October 4, 2026

AI-Infra-Guard for auditing AI infrastructure, agents and skills from one platform

AI-Infra-Guard for auditing AI infrastructure, agents and skills from one platform

Tencent's AI-Infra-Guard consolidates AI red teaming into a single self-hosted platform combining ClawScan, Agent Scan, MCP Server scanning, skill auditing and jailbreak evaluation for authorized security teams.

ToolTencent/AI-Infra-Guard — full-stack AI red teaming platform by Tencent Zhuque Lab, Apache-2.0, ~6.4k stars
CategoryAI security scanning / red teaming platform (Python)
Primary UseSelf-examination of AI ecosystems: infra CVE scanning, Agent Scan, MCP server and agent skill auditing, jailbreak evaluation
Safe UseDeployed internally by enterprises or individuals for authorized assessment of their own AI stacks; the README explicitly warns it has no authentication and must not be exposed to public networks
Telemetry NoteServer-side platform generating structured scan reports (result.json); defenders can monitor its docker-compose footprint, API endpoints like /api/v1/relay/check/stream, and CI/CD integration points

AI-Infra-Guard, published by Tencent's Zhuque Lab under Apache-2.0, is a self-hosted AI red teaming platform that tries to unify what is currently a fragmented tooling landscape. Rather than running separate scanners for AI infrastructure components, agent frameworks, MCP servers and model APIs, it bundles ClawScan (OpenClaw security scanning), Agent Scan, infrastructure vulnerability scanning, MCP Server and agent skills scanning, and jailbreak evaluation behind one web interface. At roughly 6,400 stars and an active release cadence — v4.6.2 landed in September 2026 — it has clearly struck a nerve with security teams who now own AI attack surface they never asked for.

Architecturally, the project ships as a Docker-based stack. The recommended path clones the repo and runs docker-compose -f docker-compose.images.yml up -d, pulling prebuilt images and exposing the web UI at http://localhost:8088. A one-click docker.sh install script and a build-from-source path using plain docker-compose up -d are also documented. The README is unusually candid about a deployment constraint that matters operationally: the platform currently lacks any authentication mechanism and is positioned strictly for internal enterprise or individual use — putting it on a public network is explicitly discouraged. That single sentence tells you the maintainers understand their audience and their threat model.

What gives the platform its teeth is the vulnerability library. The v4.6.0 release expanded coverage to 146 AI components and over 2,000 CVE rules, and v4.6.2 added another 155 rules across 40-plus components including LangFlow, n8n, PraisonAI, vLLM, llama-cpp and MLflow. This is the layer that turns the tool from a research curiosity into something a practitioner can point at a GPU cluster or an orchestration deployment and get a concrete, version-aware risk picture. The component list itself is a useful map of what an AI stack actually consists of in 2026 — inference engines, workflow builders, agent frameworks — and each is a real product with real CVE history.

The skill-scanning side is arguably the most interesting engineering. aig-skill-scan is installable as a standalone pip install aig-skill-scan CLI designed to slot into enterprise CI/CD pipelines, taking a --repo path, an LLM model selector like -m deepseek-v4-flash, a --language flag and an -o result.json output. Its classification scheme aligns with the published SkillTrustBench taxonomy T01–T09, which covers nine risk categories across five layers: instruction and memory risks such as T01 skill instruction hijacking and T02 memory poisoning; code execution risks like T03 remote payload download and T04 embedded malicious code; privilege risks (T05 escalation, T06 persistence); toolchain risks (T07 tool hijacking, T08 insecure dependencies); and T09 insecure coding practices.

The README also publishes benchmark numbers for aig-skill-scan against SkillTrustBench, and they are worth reading critically rather than accepting at face value. With Claude Opus 4.6 the scanner reports an F1 of 0.9848 and recall of 0.9974 at a 0.0663 false positive rate; Gemini 3.5 Flash trades recall (0.9641) for much better precision (0.9947) and a notably low FPR of 0.0120. The fact that detection quality depends on the backing model is the honest takeaway — this is an LLM-assisted analyzer, not a deterministic signature matcher, and the model you point it at shapes your noise profile. The v4.6.2 changelog entry about reducing evidence-free false positives in Skill-Scan suggests the team is actively fighting that battle.

Recent releases show a pattern of shipping both detection depth and defensive hardening. v4.5.2 added .pyc bytecode bypass detection and charset smuggling defense to Skill-Scan, plus RCE prevention via tool whitelisting in MCP-Scan's dynamic mode — an acknowledgement that a dynamic scanner itself is an execution surface. v4.6.0 introduced LLM API poisoning detection, a multi-probe black-box audit for model substitution and backdoor risks, alongside an Agent-Scan v5.0.0 mutation engine refactor and new agent loss-of-control benchmarks (FORGE-Bench, RogueHandoff-20). v4.5.1 brought four multi-turn jailbreak attack models into the evaluation module: Many-Shot, PAIR, GOAT and ActorAttack, plus web-exfiltration detection in Agent-Scan and additional MCP-Scan rules.

The API relay checker is a separate service worth noting for defenders. Docker deployment exposes GET /api/v1/relay/models and POST /api/v1/relay/check/stream, with Swagger docs at http://127.0.0.1:8088/api-checker/docs; running it from source involves a Python venv for services/api_checker plus a Go build of the unified CLI via go build -o ai-infra-guard ./cmd/cli/main.go. v4.6.1 expanded its model fingerprint coverage to include Gemini 2.5/3.1, Gemma 2/3/4 and GLM-5.3 variants — the checker's purpose is verifying that the model behind a relay endpoint is what it claims to be, which addresses a real supply-chain concern for anyone routing production traffic through third-party LLM proxies.

Integration reach is another signal of maturity. The platform publishes standalone skill-scan, mcp-scan and agent-scan CLI directories within the repo, is callable from OpenClaw chat via a clawhub install aig-scanner skill (configured through an AIG_BASE_URL environment variable), and appears on ClawHub with EdgeOne ClawScan, EdgeOne Skill Scanner and AIG Scanner badges. It was presented in the Black Hat EU 2025 Arsenal, is listed in DeepSeek's awesome-deepseek-integration, and carries a Trendshift badge — a mix of community and industry validation that few AI security tools have accumulated this quickly.

There is also a commercial angle that purchasers should understand before deploying. An online Pro version exists at aigsec.ai behind an invitation-code gate, prioritized for active contributors, and Tencent operates a related skill-market ecosystem. The core platform remains open source, but the existence of a hosted tier means feature differentiation between community and Pro is something to evaluate rather than assume. The README's CHANGELOG.md, multilingual translations (Chinese, Japanese, Spanish, German, French, Korean, Portuguese, Russian), and a user feedback survey with incentives all point to a project investing heavily in adoption rather than abandoning ship after a launch spike.

For an authorized practitioner, the workflow this tool enables is straightforward self-examination: deploy the Docker stack on an internal host, point the infra scanner at your AI deployment inventory, wire aig-skill-scan into the pipeline that vets agent skills before they ship, and run the jailbreak evaluation module against models you operate before adversarial users do it for you. Everything the README documents is oriented toward auditing your own estate — there is no operational tooling here for compromising third-party systems. Respecting the maintainers' own warning matters: bind the platform to internal networks, keep it behind your VPN, and treat its unauthenticated UI as the liability it is.

The telemetry and detection perspective is relevant both ways. As a defender, knowing that AI-Infra-Guard exists means knowing that internal ad-hoc deployments of it may appear in your environment as Docker containers answering on port 8088 — an unauthenticated management surface you should inventory. Conversely, its own findings — CVE-matched components, hijacked skills, poisoned memory, non-whitelisted tool calls — are exactly the artifact classes your detection engineering should be consuming from it, particularly the T01–T09 taxonomy, which is a reasonable shared vocabulary for alerting on agent misbehavior regardless of which scanner produced the signal.

Caveats to weigh: the project moves fast enough that CVE rule quality varies (v4.6.1 corrected mislabeled CVE product names, and v4.6.0+ fixed MCP-Scan runs that produced empty or incomplete results), and LLM-backed analysis inherits the failure modes of its underlying model. The documentation lives at tencent.github.io/AI-Infra-Guard, and the contribution norms — issues, pull requests, discussions — feed invitation codes for the Pro tier, which is a sensible community-building loop. As a consolidated, actively maintained, honestly scoped platform for auditing AI infrastructure you are authorized to test, AI-Infra-Guard is currently one of the strongest open-source options available.

Official project repository for Tencent/AI-Infra-Guard.
Download Tool

Educational analysis for authorized security professionals. Use only in controlled, authorized environments.

Share articleFacebookXLinkedIn

Continue exploring

Browse all articles →

0 comentários:

Post a Comment

Note: Only a member of this blog may post a comment.