{
  "markdown": "# Web Scraping Skill\n\nIntelligent web scraping with automatic strategy selection and TypeScript-first Apify Actor development.\n\n## Overview\n\nThis skill provides:\n- **Adaptive reconnaissance** - Phases 0-5 with quality gates that skip unnecessary work (curl first, browser only if needed)\n- **Framework-aware detection** - Identifies site framework before searching, skips irrelevant patterns\n- **Validated findings** - Every claimed selector/path/API is tested before reporting\n- **Self-critiquing reports** - Intelligence reports include gap analysis and staleness warnings\n- **Iterative implementation** - Starts simple, adds complexity only if needed\n- **Production-ready guidance** - TypeScript-first Apify Actor development\n\n## Installation\n\nAdd this skill to Claude Code by placing this directory in the skills folder.\n\n## Quick Start\n\n### Scenario 1: Scrape a Website\n\n```\nUser: \"Scrape https://example.com\"\n\nClaude will automatically:\n1. Phase 0: curl raw HTML — detect framework, search for data points, check sitemaps\n2. QUALITY GATE: All data in HTML? → Skip browser, go to validation\n3. Phase 1: Launch stealth browser (only if needed) — capture traffic, rendered DOM\n4. Phase 2: Deep scan (only for missing data) — test interactions, sniff APIs\n5. Phase 3: Validate every finding — test selectors, replay APIs, confirm paths\n6. Phase 4: Protection testing (only if signals detected or user requested)\n7. Phase 5: Generate intelligence report with self-critique\n8. Implement recommended approach iteratively\n9. Test with small batch, then scale\n```\n\n### Scenario 2: Create Apify Actor\n\n```\nUser: \"Make this an Apify Actor\"\n\nClaude will:\n1. Recommend TypeScript (strongly)\n2. Guide through `apify create` command\n3. Help choose appropriate template (Cheerio vs Playwright)\n4. Port scraping logic to Actor format\n5. Configure input schema\n6. Test and deploy\n```\n\n## Directory Structure\n\n```\nweb-scraping/\n├── SKILL.md                    # Main entry point (proactive workflow)\n├── workflows/                  # Implementation patterns\n│   ├── reconnaissance.md       # Phase 1 interactive reconnaissance (CRITICAL)\n│   ├── implementation.md       # Phase 4 iterative implementation\n│   └── productionization.md    # Phase 5 Actor creation\n├── strategies/                 # Deep-dive guides\n│   ├── framework-signatures.md # Framework detection lookup tables\n│   ├── cheerio-vs-browser-test.md # Cheerio vs Browser decision + early exit\n│   ├── proxy-escalation.md    # Protection testing skip/run conditions\n│   ├── traffic-interception.md # MITM proxy traffic capture\n│   ├── sitemap-discovery.md   # 60x faster URL discovery\n│   ├── api-discovery.md       # 10-100x faster than scraping\n│   ├── dom-scraping.md        # DevTools bridge + humanizer\n│   ├── cheerio-scraping.md    # HTTP-only (5x faster)\n│   ├── hybrid-approaches.md   # Combining strategies\n│   ├── anti-blocking.md       # Multi-layer anti-detection\n│   └── session-workflows.md   # Session recording, HAR, replay\n├── examples/                   # Runnable code\n│   ├── traffic-interception-basic.js\n│   ├── sitemap-basic.js\n│   ├── api-scraper.js\n│   ├── hybrid-sitemap-api.js\n│   └── iterative-fallback.js\n├── reference/                  # Quick lookup\n│   ├── report-schema.md       # Intelligence report format (Sections 1-7)\n│   ├── proxy-tool-reference.md # Proxy-MCP tools (80+)\n│   ├── regex-patterns.md\n│   ├── fingerprint-patterns.md\n│   └── anti-patterns.md\n├── apify/                      # Production deployment\n│   ├── typescript-first.md    # Why TypeScript\n│   ├── cli-workflow.md        # apify create (CRITICAL)\n│   ├── templates/             # TypeScript boilerplate\n│   └── examples/              # Working actors\n└── README.md                   # This file\n```\n\n## Best Practices Applied\n\nThis skill follows Anthropic's official best practices for skill development:\n\n### 1. Progressive Disclosure Architecture ✓\n\n**Pattern**: Three-level loading system to manage context efficiently\n- **Level 1**: YAML frontmatter (~85 tokens) - Always loaded\n- **Level 2**: Main SKILL.md (~356 lines) - Loaded when skill invoked\n- **Level 3**: Subdirectories - Loaded on-demand as needed\n\n**Result**: 70-80% token reduction vs monolithic documentation\n\n**Source**: [skill-creator/SKILL.md](https://github.com/anthropics/skills/blob/main/skill-creator/SKILL.md#progressive-disclosure-design-principle)\n\n### 2. Imperative/Infinitive Form Writing Style ✓\n\n**Pattern**: Write instructions using verb-first commands, not second-person language\n\n**Examples**:\n- ✅ \"Load this workflow when user requests\"\n- ✅ \"Check for sitemaps automatically\"\n- ❌ \"You should load this workflow\"\n- ❌ \"You need to check for sitemaps\"\n\n**Exception**: Second-person is acceptable in user-facing prompts, code comments, and tutorial examples\n\n**Source**: [skill-creator/SKILL.md](https://github.com/anthropics/skills/blob/main/skill-creator/SKILL.md#update-skillmd)\n\n### 3. Clear YAML Frontmatter ✓\n\n**Pattern**: Concise, specific name and description that determine when Claude invokes the skill\n\n**Applied**:\n- `name: web-scraping` - Clear, hyphen-case identifier\n- `description:` - Specific about activation triggers and capabilities (189 chars, optimized from 244)\n\n**Source**: [agent_skills_spec.md](https://github.com/anthropics/skills/blob/main/agent_skills_spec.md#yaml-frontmatter)\n\n### 4. Lean SKILL.md with Reference Files ✓\n\n**Pattern**: Keep only essential procedural instructions in SKILL.md; move detailed information to subdirectories\n\n**Applied**:\n- SKILL.md: Core 4-phase workflow (~356 lines)\n- `workflows/`: Detailed implementation patterns\n- `strategies/`: Deep-dive guides\n- `examples/`: Runnable code\n- `reference/`: Quick lookup patterns\n- `apify/`: Production deployment guides\n\n**Source**: [skill-creator/SKILL.md](https://github.com/anthropics/skills/blob/main/skill-creator/SKILL.md#references-references)\n\n### 5. Scripts, References, and Assets Organization ✓\n\n**Pattern**: Separate executable code, documentation, and output resources\n\n**Applied**:\n- `examples/` - Executable JavaScript learning examples (like scripts/)\n- `workflows/`, `strategies/`, `reference/`, `apify/` - Documentation loaded as needed (like references/)\n- `apify/templates/`, `apify/examples/` - Boilerplate code and templates (like assets/)\n\n**Source**: [skill-creator/SKILL.md](https://github.com/anthropics/skills/blob/main/skill-creator/SKILL.md#bundled-resources-optional)\n\n### 6. Purpose-Driven Skill Scope ✓\n\n**Pattern**: Create focused skills for specific purposes rather than one skill that does everything\n\n**Applied**: This skill focuses specifically on web scraping and Apify Actor development, not general web development\n\n**Source**: [Anthropic Skills Best Practices](https://www.anthropic.com/news/skills)\n\n### 7. Objective, Instructional Language ✓\n\n**Pattern**: Use clear, technical language focused on \"what\" and \"how\" rather than persuasive or promotional tone\n\n**Applied**: Direct technical guidance throughout (\"Check for sitemaps\", \"Implement iteratively\") vs. marketing language\n\n**Source**: [skill-creator/SKILL.md](https://github.com/anthropics/skills/blob/main/skill-creator/SKILL.md#update-skillmd)\n\n## Key Features\n\n### 1. Adaptive Reconnaissance (Phases 0-5)\n\nQuality-gated workflow that skips unnecessary phases:\n- **Phase 0**: curl-based assessment — detect framework, search for data, check protections\n- **Phase 1**: Browser only if needed — stealth Chrome, traffic capture, rendered DOM\n- **Phase 2**: Deep scan only for missing data — targeted interactions, framework-aware API sniffing\n- **Phase 3**: Validate every finding — test selectors, replay APIs, confirm JSON paths\n- **Phase 4**: Protection testing only if signals warrant — conditional escalation\n- **Phase 5**: Self-critiquing report — gaps, assumptions, staleness warnings\n\n### 2. Framework-Aware Detection\n\nUses `strategies/framework-signatures.md` lookup tables:\n- Response headers → framework identification\n- HTML signatures → data location mapping\n- Known major sites → direct strategy (e.g., Amazon: custom SSR, no JSON-LD)\n- Detect first, then search only relevant patterns\n\n### 3. Validated Intelligence Reports\n\nReports follow `reference/report-schema.md` with:\n- `Validated?` column for every extraction strategy (YES / PARTIAL / NO)\n- Self-Critique section: gaps, skipped steps, assumptions, staleness risk\n- Targeted re-investigation for fixable gaps\n\n### 4. Iterative Implementation (Phase 4)\n\n- Start with simplest approach\n- Test small batch (5-10 items)\n- Scale or fallback based on results\n- Add robustness last\n\n### 5. TypeScript-First Apify (Phase 5)\n\nFor production actors:\n- **Strongly recommend** TypeScript\n- **Always use** `apify create` command\n- **Choose template** based on site type (Cheerio for static, Playwright for JS-heavy)\n- Type-safe input/output\n\n## Example Workflows\n\n### Workflow 1: Unknown Site\n\n```\n1. User: \"Scrape example.com\"\n2. Phase 0: curl raw HTML → detect Next.js (__NEXT_DATA__), find product data in JSON\n3. GATE A: All data in __NEXT_DATA__? → YES → Skip browser\n4. Phase 3: Validate JSON paths resolve to expected values\n5. Phase 5: Generate report with self-critique\n6. Result: No browser needed, Cheerio + JSON parsing sufficient\n```\n\n### Workflow 1b: Site Needing Browser\n\n```\n1. User: \"Scrape protected-shop.com\"\n2. Phase 0: curl returns 403 → protection detected, no data in HTML\n3. GATE A: NO → Continue to Phase 1\n4. Phase 1: Stealth browser loads page, traffic reveals API endpoint\n5. GATE B: All data covered via API → Skip Phase 2\n6. Phase 3: Replay API request, validate response structure\n7. Phase 4: Protection testing (403 was detected) → stealth browser + proxy needed\n8. Phase 5: Report + self-critique\n9. Implements with discovered API + upstream proxies\n10. Tests with 10 items, scales to full dataset\n```\n\n### Workflow 2: Make it an Actor\n\n```\n1. User: \"Make this an Apify Actor\"\n2. Claude loads apify/ module\n3. Recommends TypeScript? (Yes)\n4. Guides through: apify create\n5. Analyzes site: Static HTML → Selects Cheerio template\n6. Ports scraping logic to TypeScript\n7. Adds input schema\n8. Tests: apify run\n9. Deploys: apify push\n10. Result: Production-ready actor\n```\n\n## Performance Benefits\n\n| Approach | Time (1000 pages) | vs Crawling |\n|----------|-------------------|-------------|\n| Sitemap + API | 5 minutes | 60x faster |\n| Sitemap + Playwright | 20 minutes | 15x faster |\n| API only | 8 minutes | 40x faster |\n| Playwright crawl | 45 minutes | Baseline |\n\n## Best Practices Summary\n\n### Reconnaissance (Phases 0-5)\n- Start with curl (Phase 0) before launching browser\n- Detect framework first, then search relevant patterns only\n- Quality gates skip phases when data is sufficient\n- Validate every selector/path/API before reporting\n- Self-critique: check for gaps, assumptions, staleness\n- Protection testing only when signals warrant it\n\n### Implementation Phase (Phase 4)\n- Start simple (traffic interception → sitemap → API → DOM scraping)\n- Test small batch first\n- Handle errors gracefully\n- Respect rate limits\n\n### Production Phase (Phase 5)\n- Use TypeScript for Apify Actors\n- Always use `apify create` command\n- Choose template based on Phase 1 findings (Cheerio vs Playwright)\n- Test locally with `apify run`\n- Deploy with `apify push`\n\n## Troubleshooting\n\n### \"No URLs found in sitemap\"\n→ See `strategies/sitemap-discovery.md` troubleshooting section\n\n### \"API requires authentication\"\n→ See `strategies/api-discovery.md` authentication section\n\n### \"DOM scraping too slow\"\n→ See `strategies/dom-scraping.md` and consider API discovered via traffic capture\n\n### \"Actor deployment fails\"\n→ See `apify/cli-workflow.md` common issues section\n\n## Resources\n\n- **Main skill**: Read `SKILL.md` for complete workflow\n- **Workflows**: Implementation patterns in `workflows/`\n- **Strategies**: Browse `strategies/` for detailed guides\n- **Examples**: Run code in `examples/` directory\n- **Reference**: Quick lookups in `reference/`\n- **Apify**: Production deployment in `apify/`\n\n## Philosophy\n\n**Intelligence first, implementation second!**\n\nThis skill prioritizes:\n1. **Reconnaissance** - Understand before coding (APIs > Sitemaps > Scraping)\n2. **Speed** - Fastest approach that works (API 10-100x faster than HTML)\n3. **Reliability** - Structured data > HTML parsing\n4. **Maintainability** - TypeScript, proper tooling\n5. **Best practices** - Industry standards\n\n## Version\n\n**5.0.0** - Traffic-interception-first scraping:\n- **NEW**: Proxy-MCP integration (MITM traffic interception + stealth browser + humanizer)\n- **NEW**: Automatic API discovery via traffic capture\n- **NEW**: Multi-layer anti-detection (stealth mode, humanizer, upstream proxies, TLS spoofing)\n- **NEW**: Session recording and HAR export/replay\n- Progressive disclosure architecture\n- Proactive strategy discovery\n- TypeScript-first Apify guidance\n- Comprehensive examples\n- Modular organization\n\n## References\n\nAll best practices sourced from official Anthropic documentation:\n- [Anthropic Skills Repository](https://github.com/anthropics/skills)\n- [Agent Skills Specification](https://github.com/anthropics/skills/blob/main/agent_skills_spec.md)\n- [skill-creator/SKILL.md](https://github.com/anthropics/skills/blob/main/skill-creator/SKILL.md)\n- [Anthropic Skills Announcement](https://www.anthropic.com/news/skills)\n\n---\n\n**Start here**: Read `SKILL.md` for the complete proactive workflow.\n",
  "bytes": 13418,
  "sha": "1a80a54e2d1cf3b3a830d5320b42feec49a237b474a93582b19aaafdc997c4e4",
  "repo_slug": "yfe404/web-scraper",
  "fonte": "repo",
  "truncated": false,
  "api": "https://api.agentalog.com/api/listings/skl_yfe404_web_scraper_web_scraping_fda6f673/readme"
}