
trawl is a self-hosted TypeScript scraping engine that fetches JS-protected pages and solves supported CAPTCHAs, positioned as an authorized-use alternative to FlareSolverr for the *arr ecosystem and AI agents.
| Tool | germondai/trawl — self-hosted scraping engine with adaptive HTTP-to-browser escalation and native CAPTCHA handling |
| Category | web scraping / anti-bot mitigation / browser automation |
| Primary Use | Authorized data retrieval from sites you own or have permission to access, drop-in FlareSolverr replacement for Prowlarr/Jackett/Sonarr, and content fetching for MCP-connected AI agents |
| Safe Use | Intended for self-hosted labs, personal media stacks, and authorized assessments or research on systems whose owners permit scraping; not for violating terms of service or third-party infrastructure |
| Telemetry Note | Its browser traffic is deliberately engineered to hide automation signals (Camoufox fingerprint patching), so defenders should focus on behavioral anomaly detection — request cadence, TLS/JA4 divergence from real Firefox, residential IP ranges, and clearance-cookie reuse patterns |
trawl describes itself as a self-hosted web scraping engine for applications and AI agents, written in TypeScript and built on the Bun runtime with an Elysia API layer. The core idea is escalation: a request begins as a plain HTTP fetch, and only climbs to a full browser solve when the target responds with a JavaScript challenge or CAPTCHA wall. That four-tier pipeline — plain HTTP fetch, cached browser session, fresh challenge solve, and an optional residential proxy — is the architectural spine of the whole project, and it explains why the authors benchmark it against FlareSolverr and Byparr rather than against generic scraper frameworks.
The second half of the architecture is session economics. Solved cookies and browser identity are persisted in Redis, so subsequent requests to a site that already accepted a session can skip the expensive fresh solve entirely. This is where the claimed performance advantage over FlareSolverr comes from: on same-machine benchmarks the authors report faster responses, with the honest caveat that results vary by site and session state. From a defender's perspective this is also the tell — reused clearance cookies tied to a stable browser identity produce request patterns that don't match organic traffic.
CAPTCHA handling is native rather than delegated to a paid solver API. The README lists support for Cloudflare Turnstile and Interstitial pages, reCAPTCHA v2 (solved via audio transcription, using Google's free STT endpoint or an optional local Whisper service), hCaptcha, GeeTest v4 slide puzzles, ALTCHA, and Friendly Captcha v1/v2. The reCAPTCHA approach is notable because the authors explicitly avoid commercial solver services, which keeps the whole stack self-contained behind docker compose.
Multi-WAF awareness is a headline feature: dedicated detection and browser flows exist for Cloudflare, Akamai Bot Manager, and Imperva/Incapsula. The engine doesn't use stock automation browsers either — it runs Camoufox, a Firefox fork patched at the C++/Juggler level to suppress the usual automation signals. This is a meaningful engineering choice over bolting a stealth plugin onto Chromium, and it's the detail that will interest anyone studying the arms race between headless browsers and bot-management vendors.
For the self-hosting community, the most immediately useful feature is the FlareSolverr-compatible /v1 endpoint. Point Prowlarr, Jackett, Sonarr, or any other *arr tool at the instance and it just works — the README gives http://localhost:8191 for same-host deployments and http://trawl:8191 for Docker Compose networks. The native POST /scrape endpoint is richer, returning tier, timings, sessionCached, and the full cookie list, which makes it a better fit for scripted pipelines that want observability into why a request succeeded or stalled.
The cleverest piece of engineering is the challenge-aware MITM proxy. The README explains a real limitation of the /v1 flow: some sites bind Cloudflare clearance to the solving browser's full connection fingerprint, so when Prowlarr re-fetches with its own HTTP client holding only the cookie and user-agent, Cloudflare re-challenges — the cookie isn't portable. TRAWL's forward proxy solves this by intercepting the client's traffic itself, escalating through the tier pipeline when it detects a challenge in the buffered response, while streaming large binaries and relaying WebSocket upgrades directly. It also honors Range/206 requests end to end.
That proxy deserves careful handling from an operational security standpoint, and the README says so explicitly. A MITM proxy that mints per-host certificates from a self-generated root CA can impersonate any host to a client that trusts that CA. The authors warn to expose it only on localhost or a private Docker network, never publicly, and the MITM_HOST=127.0.0.1 option restricts the listener to loopback on bare-metal hosts. Escalation is also capped via MITM_MAX_TIER, so an operator can, for example, set 3 to stay off residential proxies entirely — a sensible cost and exposure control.
Deployment is straightforward: clone the repo, copy .env.example to .env, run docker compose up -d to bring up the scraper plus Redis, and verify with curl http://localhost:8191/health. First boot takes 15–30 seconds while the browser pool warms up. For NAS operators, community packages exist in the TrueNAS and Unraid app catalogs, which tells you a lot about the target audience — this is a homelab tool first, maintained under AGPL-3.0 on the dev default branch, with roughly 865 stars at the time of writing.
The MCP integration is where trawl diverges from the FlareSolverr lineage. With MCP_ENABLED=true, the /mcp endpoint exposes Streamable HTTP tools that any MCP-compatible AI agent can call to read pages as Markdown, pull raw HTML, extract structured JSON records, capture viewport/full-page/element screenshots, and inspect browser diagnostics. The README is careful to note these tools load known public URLs only — the project explicitly does not provide web search or ranking, which frames the MCP surface as content retrieval rather than reconnaissance.
Observability is a first-class concern. A local dashboard at http://localhost:8191/dashboard shows persistent request history, tier outcomes, failure causes, and live updates — the kind of telemetry that matters when you're debugging why a specific indexer keeps landing in tier-3 solves. Everything shown is scoped to your own instance's traffic, and the screenshot in the docs uses illustrative data, which is a small but welcome honesty marker from the maintainers.
For defensive practitioners, trawl is worth studying even if you never run it, because it documents in one place the current state of the anti-bot bypass stack: C++-level fingerprint patching, session caching in Redis, audio-transcription CAPTCHA solving, fingerprint-bound clearance cookies, and tiered escalation through residential exits. Blue teams countering this class of tooling should assume cookie-only challenges are insufficient, that TLS and behavioral fingerprinting will outperform CAPTCHA walls, and that residential proxy traffic means IP reputation alone won't save you.
The licensing and sponsorship picture rounds out the assessment. AGPL-3.0 means anyone offering it as a service must share source, and the README's sponsor section is dominated by proxy vendors (Bright Data, NodeMaven, Thordata) with discount codes — a normal monetization pattern for this niche, but one that signals the residential-proxy tier is a real, used capability rather than a checkbox. Operate it only against properties you own, are licensed to automate, or have explicit written permission to test, and keep the MITM proxy off any interface that isn't private.
germondai/trawl.Educational analysis for authorized security professionals. Use only in controlled, authorized environments.
0 comentários:
Post a Comment
Note: Only a member of this blog may post a comment.