{
  "markdown": "<p align=\"center\">\n  <img src=\"src/assets/images/icon.png\" alt=\"Scrape-LE Logo\" width=\"96\" height=\"96\"/>\n</p>\n<h1 align=\"center\">Scrape-LE: Zero Hassle Scrapeability Checks</h1>\n<p align=\"center\">\n  <b>Load a URL in headless Chromium and see what will block your scraper — before you write it</b><br/>\n  <i>Anti-bot vendors, rate limits, robots.txt rules, login walls, console errors, screenshots</i>\n</p>\n\n<p align=\"center\">\n  <a href=\"https://marketplace.visualstudio.com/items?itemName=nolindnaidoo.scrape-le\">\n    <img src=\"https://img.shields.io/badge/Install%20from-VS%20Code-blue?style=for-the-badge&logo=visualstudiocode\" alt=\"Install from VS Code Marketplace\" />\n  </a>\n  <a href=\"https://open-vsx.org/extension/OffensiveEdge/scrape-le\">\n    <img src=\"https://img.shields.io/open-vsx/dt/OffensiveEdge/scrape-le?style=for-the-badge&label=Open%20VSX&color=blue\" alt=\"Open VSX downloads\" />\n  </a>\n  <a href=\"https://www.npmjs.com/package/scrape-le-mcp\">\n    <img src=\"https://img.shields.io/npm/v/scrape-le-mcp?style=for-the-badge&label=MCP%20server&color=blue&logo=npm\" alt=\"scrape-le-mcp on npm\" />\n  </a>\n  <a href=\"https://crates.io/crates/scrape-le\">\n    <img src=\"https://img.shields.io/crates/v/scrape-le?style=for-the-badge&label=Rust%20CLI&color=blue&logo=rust\" alt=\"scrape-le on crates.io\" />\n  </a>\n  <a href=\"https://letools.dev/tools/scrape-le\">\n    <img src=\"https://img.shields.io/badge/LE%20Tools-letools.dev-blue?style=for-the-badge\" alt=\"LE Tools\" />\n  </a>\n</p>\n\n---\n\n<p align=\"center\">\n  <img src=\"src/assets/images/demo.gif\" alt=\"Scrapeability Check Demo\" style=\"max-width: 100%; height: auto;\" />\n</p>\n\n> **Useful?** A star or rating is how other developers find it —\n> [★ GitHub](https://github.com/nolindnaidoo/scrape-le) ·\n> [★ Open VSX](https://open-vsx.org/extension/OffensiveEdge/scrape-le/reviews) ·\n> [★ Marketplace](https://marketplace.visualstudio.com/items?itemName=nolindnaidoo.scrape-le&ssr=false#review-details)\n\n## What it does\n\nRun `Scrape-LE: Check URL Scrapeability` (`Ctrl+Alt+S` / `Cmd+Alt+S`), enter a URL, and the page loads in a real headless Chromium. The report lands in the output channel: HTTP status, page title, load time, console errors, a full-page screenshot, and four detections. Works in VS Code and VS Code–based editors like Cursor and VSCodium (installable from Open VSX).\n\nOne-time setup: run `Scrape-LE: Setup Browser` to install Chromium (~130MB, into Playwright's browser cache).\n\n## Install\n\n| Where | What you get | Install |\n|---|---|---|\n| **VS Code** | The same check, in your editor, on a keystroke | [Marketplace](https://marketplace.visualstudio.com/items?itemName=nolindnaidoo.scrape-le) |\n| **Cursor, VSCodium, Windsurf** | The same extension | [Open VSX](https://open-vsx.org/extension/OffensiveEdge/scrape-le) |\n| **A terminal or a CI step** | The same run over a whole tree, with exit codes | `cargo install scrape-le` · [crates.io](https://crates.io/crates/scrape-le) |\n| **Any MCP agent, via Node** | `analyze_robots_txt` over stdio | `npx scrape-le-mcp` · [npm](https://www.npmjs.com/package/scrape-le-mcp) |\n| **Zed** | The MCP server as a context server | [add it by hand](https://zed.dev/docs/ai/mcp) *(no listing yet)* |\n\n## Use it from an AI agent\n\nThe same engine runs as an [MCP](https://modelcontextprotocol.io) server, so an agent can call it directly instead of you running a command.\n\n| Editor | How |\n|---|---|\n| **VS Code** 1.101+ | Nothing to install — the extension registers `analyze_robots_txt` with agent mode |\n| **Zed** | No listing yet — [add the MCP server by hand](https://zed.dev/docs/ai/mcp) |\n| **Claude Code** | `claude mcp add scrape-le -- npx -y scrape-le-mcp` |\n| **Cursor, Windsurf, anything else** | point it at `npx scrape-le-mcp` |\n\n```\nanalyze_robots_txt(content, path, maxResults?)\n```\n\nGiven robots.txt contents and a path, reports whether the generic (`User-agent: *`) rules permit crawling it, plus the crawl delay, disallowed patterns and any sitemaps.\n\nThe server takes content and returns data — it reads no files and makes no network requests of its own. Published as [`scrape-le-mcp`](https://www.npmjs.com/package/scrape-le-mcp) on npm and as `io.github.nolindnaidoo/scrape-le` in the [MCP registry](https://registry.modelcontextprotocol.io).\n\n<details>\n<summary><b>Configuring it by hand</b> — any host with an MCP config file</summary>\n\nMost hosts read a JSON config. Add one entry:\n\n```json\n{\n  \"mcpServers\": {\n    \"scrape-le\": {\n      \"command\": \"npx\",\n      \"args\": [\"-y\", \"scrape-le-mcp\"]\n    }\n  }\n}\n```\n\n`-y` skips the install prompt on first run. Pin a version if you would rather not track releases — `scrape-le-mcp@2.2.6`.\n\nPrefer not to go through `npx` on every launch? Install it once and point at the binary instead:\n\n```bash\nnpm install -g scrape-le-mcp\n```\n\n```json\n{\n  \"mcpServers\": {\n    \"scrape-le\": { \"command\": \"scrape-le-mcp\" }\n  }\n}\n```\n\nIt speaks MCP over stdio and needs no environment variables, no API key and no configuration of its own. To check it before wiring it into anything:\n\n```bash\necho '{\"jsonrpc\":\"2.0\",\"id\":1,\"method\":\"tools/list\"}' | npx -y scrape-le-mcp\n```\n\nThat prints the tool list and exits — if you see `analyze_robots_txt`, the server works.\n\n</details>\n\n## The CLI\n\nThe same check runs from a terminal or an agent loop: a Rust CLI in [`crate/`](crate/) of this repository, sharing one signature corpus with the extension — [`crate/signatures/`](crate/signatures/) and [`crate/fixtures/`](crate/fixtures/) — so CI fails if the two ever disagree about a URL.\n\n```bash\nscrape-le https://example.com/search   # JSON on stdout, summary on stderr\nscrape-le --input urls.txt             # a batch, streamed as it completes\nscrape-le mcp                          # the same check over MCP on stdio\n```\n\nThe exit code is the answer: **0 clear · 1 a real no · 2 the question was malformed.** ## Detections\n\n| Detection | How it works |\n|---|---|\n| Anti-bot vendors | Response headers, script sources, DOM elements, and window globals fingerprint Cloudflare (incl. Turnstile challenges), reCAPTCHA, hCaptcha, DataDome, and PerimeterX |\n| Rate limiting | `X-RateLimit-*` / `RateLimit-*` / `Retry-After` response headers, plus HTTP 429 |\n| robots.txt | Fetches `<origin>/robots.txt` and evaluates the `User-agent: *` rules against your URL with RFC 9309 semantics — grouped agents, `Allow`/`Disallow` longest-match, `*` wildcards, `$` anchors, crawl-delay, sitemaps |\n| Authentication | HTTP 401/403, login forms (password + username fields), auth keywords in page text, auth path segments in the final URL |\n\nHonest limitations: signatures are best-effort fingerprints of public integration patterns — a detected widget means the page *can* challenge you, not that it will, and a clean result is not proof a site allows scraping. Agent-specific robots.txt groups are ignored (only the `*` rules are reported). Pages get up to 5 seconds to go network-idle after load, so content rendered later than that can be missed by the page-level detections.\n\n## Commands\n\n| Command | Description |\n|---|---|\n| `Scrape-LE: Check URL Scrapeability` (`Ctrl+Alt+S` / `Cmd+Alt+S`) | Prompt for a URL and run the full check |\n| `Scrape-LE: Check Selected URL` | Run the check on the URL in the current selection (also in the right-click menu) |\n| `Scrape-LE: Setup Browser` | Install or verify the Chromium browser |\n| `Scrape-LE: Open Settings` | Open Scrape-LE settings |\n| `Scrape-LE: Help & Troubleshooting` | Built-in documentation |\n\n## Settings\n\n| Setting | Default | Description |\n|---|---|---|\n| `scrape-le.browser.timeout` | `30000` | Page-load timeout in ms (5000–120000) |\n| `scrape-le.browser.viewport.width` | `1280` | Viewport width |\n| `scrape-le.browser.viewport.height` | `720` | Viewport height |\n| `scrape-le.browser.userAgent` | `\"\"` | Custom User-Agent (empty = Chromium default) |\n| `scrape-le.retry.userAgents` | `false` | On a blocked or failed check, retry under common User-Agents and report which worked |\n| `scrape-le.screenshot.enabled` | `true` | Save a full-page screenshot per check |\n| `scrape-le.screenshot.path` | `.vscode/scrape-le` | Screenshot directory (workspace-relative or absolute) |\n| `scrape-le.screenshot.format` | `png` | `png` or `jpeg` |\n| `scrape-le.screenshot.quality` | `90` | JPEG quality 0–100 (ignored for png) |\n| `scrape-le.checkConsoleErrors` | `true` | Capture console and page errors while loading |\n| `scrape-le.detections.antiBot` | `true` | Anti-bot vendor detection |\n| `scrape-le.detections.rateLimit` | `true` | Rate-limit detection |\n| `scrape-le.detections.robotsTxt` | `true` | robots.txt fetch + evaluation |\n| `scrape-le.detections.authentication` | `true` | Authentication-wall detection |\n| `scrape-le.notificationsLevel` | `important` | `all` = every notification, `important` = warnings + errors, `silent` = errors only |\n| `scrape-le.statusBar.enabled` | `true` | Show the status bar item |\n\n## Languages\n\nTwelve languages besides English:\n\nGerman · Spanish · French · Indonesian · Italian · Japanese · Korean ·\nPortuguese (Brazil) · Russian · Ukrainian · Vietnamese · Chinese (Simplified)\n\nBoth halves are covered — the manifest (command titles, setting names and\ndescriptions) and everything shown while the extension runs (notifications,\nthe status bar, quick-picks and prompts). The extension follows VS Code's\ndisplay language, so it matches whatever the editor is already set to; no\nsetting of its own.\n\n## Privacy & security\n\n- **Network access is the feature, and it is scoped.** A check talks to exactly two things: the URL you enter (loaded in headless Chromium, which fetches that page's own resources like any browser) and that origin's `/robots.txt`. Nothing is sent anywhere else — no telemetry, no analytics.\n- **Screenshots stay local**, written to the configured path inside your workspace.\n- **The MCP server makes no network request at all** — unlike the extension, deliberately. `fetchRobotsTxt` builds a URL from an arbitrary origin, which inside an agent loop is an SSRF primitive: the caller supplying the URL is the model, not you. The server analyses robots.txt content you already fetched, and a test asserts no tool accepts a `url` argument.\n- Error notifications redact home directories and credential-shaped fragments.\n- Respect the sites you check: a scrapeability report is information, not permission.\n\n## Documentation\n\n| What | Where |\n|---|---|\n| What the tool is allowed to say — scope, output contract, refusals, non-goals | [`crate/SPEC.md`](crate/SPEC.md) |\n| How the extension is built and held together — architecture, invariants, toolchain, release | [AGENTS.md](AGENTS.md) |\n| How the CLI is built and held together | [`crate/AGENTS.md`](crate/AGENTS.md) |\n| What changed | [CHANGELOG.md](CHANGELOG.md) · [`crate/CHANGELOG.md`](crate/CHANGELOG.md) |\n| The tool's page, and the other fifteen | [letools.dev/tools/scrape-le](https://letools.dev/tools/scrape-le) |\n\n## Performance\n\n<!-- performance:start -->\n| Input | Size | Found | Time | Rate | Scan speed |\n| --- | --- | --- | --- | --- | --- |\n| Header signature scan | 2.83 MB | 20,000 | 5.19 ms | 3,852,946/sec | 544.6 MB/s |\n| robots.txt path match | 3.32 MB | 60,000 | 9.64 ms | 6,223,689/sec | 344.3 MB/s |\n\nMedian of 7 runs after warmup, on Apple M5 Pro, 24 GB RAM, Node 24.3.0. Inputs are generated\nby `scripts/benchmark.ts` rather than checked in, so the sizes above are\nexactly what was measured. Reproduce with `bun run benchmark`.\n\nThese are machine-specific and are not asserted in CI — a benchmark that gates\na build only tells you how busy the runner was.\n<!-- performance:end -->\n\n## Testing\n\n<!-- coverage:start -->\n| Metric | Coverage |\n| --- | --- |\n| Statements | 93.13% |\n| Branches | 84.44% |\n| Functions | 95.20% |\n| Lines | 94.54% |\n\n379 test cases across 29 files, plus an integration suite that runs\nin a real VS Code extension host and an end-to-end test that installs the\nbuilt `.vsix` into a clean profile.\n\nGenerated from a real run — `coverage/coverage-summary.json` and\n`coverage/test-results.json` — by `scripts/coverage-readme.js`; CI fails if\nthis section drifts. Reproduce with `bun run test:coverage`, and the case\ncount is the one vitest prints.\n<!-- coverage:end -->\n\n## More from the LE family\n\nSixteen single-purpose tools for the work in front of every model. Each ships\na Rust CLI and an MCP server. One page: **[letools.dev](https://letools.dev)**\n\n**Get it out**\n\n- **[String-LE](https://letools.dev/tools/string-le)** — Extract every string in a codebase, with its position, so a person can read them\n- **[Numbers-LE](https://letools.dev/tools/numbers-le)** — Extract every hardcoded number in a codebase, so a person can check them\n- **[Units-LE](https://letools.dev/tools/units-le)** — Extract every quantity with its unit, normalized, and refuse the ambiguous ones by name\n- **[Dates-LE](https://letools.dev/tools/dates-le)** — Extract every date and timestamp, and the exact instant each one resolves to\n- **[IDs-LE](https://letools.dev/tools/ids-le)** — Extract every UUID, ULID, NanoID, ObjectId and Snowflake, and decode the time inside\n- **[IPs-LE](https://letools.dev/tools/ips-le)** — Extract every IP address, CIDR block and MAC, normalized and classified by scope\n- **[URLs-LE](https://letools.dev/tools/urls-le)** — Extract every URL in a codebase, with its protocol and exact position\n- **[Paths-LE](https://letools.dev/tools/paths-le)** — Extract every file path in a codebase, and say whether it still points at anything\n- **[Colors-LE](https://letools.dev/tools/colors-le)** — Extract every color in a codebase, and say which ones are not in your palette\n\n**Check it**\n\n- **[Regex-LE](https://letools.dev/tools/regex-le)** — Find every regex in a codebase, and report which can be driven into catastrophic backtracking\n- **[Versions-LE](https://letools.dev/tools/versions-le)** — Find where one dependency is constrained differently across a repository's manifests\n- **[i18n-LE](https://letools.dev/tools/i18n-le)** — Identify the i18n library a project uses, then audit its catalogs by that library's rules\n- **[Scrape-LE](https://letools.dev/tools/scrape-le)** — Check whether a page is scrapeable before the scraper is written, and say when it cannot tell\n\n**Guard it**\n\n- **[Secrets-LE](https://letools.dev/tools/secrets-le)** — Find hardcoded credentials in a codebase, and never print one into the report\n- **[EnvSync-LE](https://letools.dev/tools/envsync-le)** — Compare the dotenv files in a tree, and say which keys are missing from which\n- **[Unicode-LE](https://letools.dev/tools/unicode-le)** — Find the Unicode that hides meaning — bidi controls, invisibles, homoglyphs, mixed scripts\n\nEach stands on its own: no shared crate, no published core. Where two of them\nagree, it is because the same answer was right twice.\n\n**Contact** — [nolindnaidoo.com](https://nolindnaidoo.com) · [GitHub](https://github.com/nolindnaidoo) · [LinkedIn](https://www.linkedin.com/in/nolindnaidoo/)\n\n## Also by nolindnaidoo\n\n**Rust** — pixelcoords and pixelactions are one loop: pixelcoords answers\n*where*, pixelactions *acts* there. Their own tools, their own voice — not\npart of the LE family.\n\n- **[pixelcoords](https://github.com/nolindnaidoo/pixelcoords)** — Freeze your screen, mark regions, get pixel-exact coordinates and crops\n  [pixelcoords.dev](https://pixelcoords.dev) · [crates.io](https://crates.io/crates/pixelcoords) · [docs.rs](https://docs.rs/pixelcoords)\n- **[pixelactions](https://github.com/nolindnaidoo/pixelactions)** — Consume human-verified coordinates, perform the interaction, confirm it landed\n  [pixelactions.dev](https://pixelactions.dev) · [crates.io](https://crates.io/crates/pixelactions) · [docs.rs](https://docs.rs/pixelactions)\n\n## License\n\nMIT © [nolindnaidoo](https://github.com/nolindnaidoo)\n",
  "bytes": 15810,
  "sha": "e70a034326c5a6b43e4bbd43f64a0962e29f3a5771108f5a2f0c4f1110cbe09a",
  "repo_slug": "nolindnaidoo/scrape-le",
  "fonte": "repo",
  "truncated": false,
  "api": "https://agentalog.com/api/listings/mcp_io_github_nolindnaidoo_scrape_le_110a1ea1/readme"
}