{
  "markdown": "This file is 184 lines long; read all of them.\n\n<p align=\"center\">\n\t<img src=\"assets/zyte-logo.png\" alt=\"Zyte\" width=\"180\">\n</p>\n\n<h1 align=\"center\">Zyte Agentic Web Data</h1>\n\n<p align=\"center\">\n\tFrom a plain-English prompt to a working Scrapy spider.\n</p>\n\n<p align=\"center\">\n\t<a href=\"https://github.com/zytedata/skills/releases/tag/0.4.0\">\n\t\t<img src=\"https://img.shields.io/badge/version-0.4.0-blue\" alt=\"Version 0.4.0\">\n\t</a>\n\t<a href=\"https://github.com/zytedata/skills/blob/main/LICENSE.md\">\n\t\t<img src=\"https://img.shields.io/badge/license-Zyte%20EULA-b02cce\" alt=\"Zyte EULA\">\n\t</a>\n\t<a href=\"https://skills.sh/zytedata/skills\">\n\t\t<img src=\"https://skills.sh/b/zytedata/skills\" alt=\"Installs\">\n\t</a>\n\t<a href=\"https://github.com/zytedata/skills\">\n\t\t<img src=\"https://img.shields.io/github/stars/zytedata/skills?style=social\" alt=\"GitHub stars\">\n\t</a>\n</p>\n\n---\n\n> Using a specific coding agent? See [Zyte Coding Agent Add-Ons](https://docs.zyte.com/ai-code.html) for alternatives.\n\n## Install\n\nThis is an [agent plugin](https://agent-plugins.org/): a portable, vendor-neutral package of skills that any agent implementing the standard can load. Install it however your agent installs plugins, or use the [`skills`](https://www.skills.sh) CLI, which sets it up for every supported agent on your system at once:\n\n```bash\nnpx skills add zytedata/skills\n```\n\nSee [skills.sh](https://www.skills.sh) for the list of agents it supports and how to enable skills in each.\n\nAny agent that reads agent plugins or [Agent Skills](https://agentskills.io) can use this plugin, including editors such as Cursor. For Claude Code, Codex CLI and GitHub Copilot CLI we publish [dedicated packages](https://docs.zyte.com/ai-code.html) tuned to each of them; prefer those if you use one.\n\nThe plugin declares the [Zyte MCP](https://docs.zyte.com/zyte-web-data/mcp.html) server in `mcp.json`, for agents that read MCP servers from plugins; sign in to it once, in the browser. Elsewhere, add it the way your agent documents MCP servers, as an HTTP server named `zyte` at `https://mcp.zyte.com/v1/mcp`.\n\n---\n\n## What it does\n\nThis is Zyte's official [agent plugin](https://agent-plugins.org/) that generates production-ready [Scrapy](https://scrapy.org) spiders with [web-poet](https://web-poet.readthedocs.io) page objects from a plain-English prompt. Give it a URL and describe what you want to extract. It handles site exploration, schema discovery, code generation, and smoke testing: no boilerplate, no manual selector hunting.\n\nThe plugin explores the target site, discovers available fields, and presents a schema for your approval before generating a single line of code. After you confirm the schema, it creates a Scrapy project with all dependencies configured, generates web-poet page objects and test fixtures, wires up the spider, and runs a smoke test to verify that extraction is working before handing the project back to you.\n\nOptionally, use `/zyte` to deploy directly to [Scrapy Cloud](https://www.zyte.com/scrapy-cloud/) for scheduled runs, job history, and monitoring, which the Zyte MCP handles once the spider is deployed. A [free tier is available](https://docs.zyte.com/scrapy-cloud/pricing.md).\n\n---\n\n## Use cases\n\nThe `/scrape` skill works on any website with repeating structured content: detail pages linked from a listing or category page. Examples from the skill:\n\n- Product catalogs\n- Job listings\n- Recipes\n\n---\n\n## How does it work?\n\nThe `/scrape` skill orchestrates two stages automatically:\n\n```\n1. Plan and validate the scrape     →  /scrape-plan\n2. Build the project and spider     →  /scrapy-extra\n```\n\nEach stage feeds directly into the next. When the pipeline completes, you have a runnable spider and a passing test suite:\n\n```bash\nuv run scrapy crawl <spider_name>\nuv run pytest fixtures/\n```\n\n---\n\n## Skills\n\n### Orchestration\n\n| Skill | Description |\n|---|---|\n| `scrape` | End-to-end web scraping workflow — from URL to working spider with web-poet page objects |\n\n### Pipeline stages (called automatically by `/scrape`)\n\n| Skill | Description |\n|---|---|\n| `scrape-plan` | Plan the scrape and author a validated extraction spec: discover fields, download diverse pages, compare HTML variants, optional browser review |\n| `scrape-analyze-page` | Extract all available fields with values from a detail page |\n| `scrapy-extra` | Hands-on Scrapy coding: write/debug spiders, web-poet page objects, and projects; configure scrapy-poet and scrapy-zyte-api |\n\n### Zyte APIs\n\n| Skill | Description |\n|---|---|\n| `zyte` | Zyte work the Zyte MCP does not cover: set up your Zyte account and credentials; deploy projects to [Scrapy Cloud](https://www.zyte.com/scrapy-cloud/) and export all items, logs or requests of a job; look up [Zyte API](https://www.zyte.com/zyte-api/) plan pricing; and answer how-to and documentation questions about Zyte from the official docs |\n\n---\n\n## Prerequisites\n\n- An AI coding agent that supports [agent plugins](https://agent-plugins.org/) or [Agent Skills](https://agentskills.io)\n- [`uv`](https://docs.astral.sh/uv/) — used to create and manage the Scrapy project\n- A [Zyte](https://www.zyte.com/) account, signed in once through the [Zyte MCP](https://docs.zyte.com/zyte-web-data/mcp.html) (see [Install](#install)) — used for Scrapy Cloud jobs, Zyte API usage stats and per-website prices\n\nProject dependencies (scrapy, scrapy-poet, scrapy-zyte-api, web-poet, extruct, price-parser, pytest) are installed automatically by the skills.\n\n---\n\n## Quickstart\n\nAny scraping prompt triggers the skill automatically. For example:\n\n```\nScrape books.toscrape.com\n```\n\nThe plugin walks you through schema approval interactively, then generates a complete, tested Scrapy project.\n\n---\n\n## Update\n\nTo update:\n\n```bash\nnpx skills update zytedata/skills\n```\n\n---\n\n## Evaluation\n\nWe automatically evaluate skills and track both wall time and cost. We measure and aim to improve these metrics over time.\n\n---\n\n## Feedback\n\nIf you find any issue — such as prompts that did not work as expected, or that caused excessive wall time or cost — please [open a GitHub issue](https://github.com/zytedata/skills/issues).\n\nProvide as much detail as possible to help us reproduce the issue. You are welcome to anonymize target websites or other data.\n\n---\n\n## Frequently asked questions\n\n### Is a Zyte account required?\n\nNo. The generated spider is a standard Scrapy project that runs locally with `uv`. A Zyte account is required only if you want to deploy to [Scrapy Cloud](https://www.zyte.com/scrapy-cloud/) or use [Zyte API](https://www.zyte.com/zyte-api/) to access sites that block standard scrapers. If you want to use Zyte API, you'll need an account to generate an API key.\n\n### Does it handle JavaScript-rendered pages?\n\nThe generated project includes `scrapy-zyte-api` as a dependency. Enabling headless browser rendering requires a [Zyte API](https://www.zyte.com/zyte-api/) key. The `/zyte` skill guides you through setting up your credentials.\n\n### What Python libraries does the generated project use?\n\nThe project template includes `scrapy`, `scrapy-poet`, `scrapy-zyte-api`, `web-poet`, `extruct`, `price-parser`, and `pytest`. All dependencies are installed automatically via `uv sync`.\n\n### Can the generated spider run without my AI coding agent?\n\nYes. The plugin generates a standard Scrapy project. Run it directly with:\n\n```bash\nuv run scrapy crawl <spider_name>\n```\n\nYou can extend, modify, and deploy it independently of your AI coding agent.\n\n---\n\n## License\n\nSee [LICENSE.md](LICENSE.md) for the Zyte End User License Agreement.\n",
  "bytes": 7587,
  "sha": "c996efc1d1709d0ae6ed5e87d24ae1409af6cb259ef0e9802389c700cc40541d",
  "repo_slug": "zytedata/skills",
  "fonte": "repo",
  "truncated": false,
  "api": "https://api.agentalog.com/api/listings/skl_zytedata_skills_scrape_ensure_project_2a18540b/readme"
}