{
  "markdown": "<p align=\"center\">\n  <img src=\"docs/assets/logo.svg\" alt=\"pSEO Lint: audit your pSEO site by template, not by URL\" width=\"720\" />\n</p>\n\n<p align=\"center\">\n  <a href=\"https://www.npmjs.com/package/pseolint\"><img src=\"https://img.shields.io/npm/v/pseolint?color=cb3837&logo=npm\" alt=\"npm\" /></a>\n  <a href=\"https://www.npmjs.com/package/pseolint\"><img src=\"https://img.shields.io/npm/dm/pseolint?color=cb3837\" alt=\"Downloads\" /></a>\n  <a href=\"./LICENSE\"><img src=\"https://img.shields.io/npm/l/pseolint?color=blue\" alt=\"License\" /></a>\n  <a href=\"https://www.npmjs.com/package/pseolint\"><img src=\"https://img.shields.io/node/v/pseolint?color=339933&logo=node.js\" alt=\"Node\" /></a>\n  <a href=\"https://github.com/ouranos-labs/pseolint\"><img src=\"https://img.shields.io/github/stars/ouranos-labs/pseolint?style=social\" alt=\"GitHub stars\" /></a>\n  <a href=\"https://glama.ai/mcp/servers/ouranos-labs/pseolint\"><img src=\"https://glama.ai/mcp/servers/ouranos-labs/pseolint/badges/score.svg\" alt=\"ouranos-labs/pseolint MCP server\" /></a>\n  <a href=\"https://pseolint.dev/leaderboard\"><img src=\"https://pseolint.dev/api/badge/pseolint.dev\" alt=\"pseolint.dev dogfood\" /></a>\n</p>\n\n<p align=\"center\">\n  <strong><a href=\"https://pseolint.dev/methodology\">Methodology</a></strong> ·\n  <strong><a href=\"https://pseolint.dev/leaderboard\">Leaderboard</a></strong> ·\n  <strong><a href=\"https://github.com/ouranos-labs/pseolint/issues/new\">Report a bug</a></strong> ·\n  <strong><a href=\"skills/README.md\">Skills for agents</a></strong>\n</p>\n\n<p align=\"center\">\n  <img src=\"docs/assets/demo.svg\" alt=\"pseolint auditing pseolint.dev: verdict READY, all four categories graded A, with a per-template breakdown of /rules/:slug, /tools/:slug, and long-tail pages\" width=\"820\" />\n</p>\n\nThe only tool purpose-built for **programmatic SEO compliance**. It shifts the unit of analysis from URL to template: point it at a 10,000-URL pSEO directory and pseolint identifies the template clusters (e.g. `/listing/:slug`, `/category/:slug`), samples K pages from each, and produces a per-template verdict + variance metric. **Fix one template, fix N pages.**\n\n```bash\nnpx pseolint http://localhost:3000\n```\n\n<details>\n<summary><strong>Table of contents</strong></summary>\n\n- [Why this exists](#why-this-exists)\n- [How pseolint differs](#how-pseolint-differs)\n- [Quick Start](#quick-start)\n- [What It Checks](#what-it-checks): the 65 rules\n- [CLI Options](#cli-options)\n- [GitHub Action](#github-action)\n- [Fix rail: from audit to pull request](#fix-rail--from-audit-to-pull-request)\n- [Skills for Claude & coding agents](#skills-for-claude--coding-agents-new)\n- [Roadmap](#roadmap)\n- [Contributing](#contributing)\n\n</details>\n\n## Skills for Claude & coding agents (new)\n\nDesign pages that pass *before* you crawl them. The [`skills/`](skills/) suite gives\nClaude (and any agent that supports the `skills` / Claude Code plugin format)\nprogrammatic-SEO and **answer-engine optimization (AEO / GEO)** guidance where\nevery recommendation is bound to a runnable pseolint rule:\n\n```bash\nnpx skills add ouranos-labs/pseolint --skill pseolint aeo\n```\n\n- **`pseolint`**: full-lifecycle programmatic SEO: design → build → audit → fix → gate.\n- **`aeo`**: get cited in AI Overviews / ChatGPT / Perplexity, not just ranked.\n\nUnlike prose checklists, these have teeth: the design-time advice ends in\n`npx pseolint` pass/fail. See [`skills/README.md`](skills/README.md).\n\n## Why this exists\n\nProgrammatic SEO works, when it works. The gap between \"1,000 indexed pages\" and \"1,000 pages that survive a SpamBrain pass\" is where most pSEO sites die. The Helpful Content Update made that gap permanent.\n\nExisting SEO tools (Screaming Frog, Sitebulb, Ahrefs Site Audit) were built for editorially-curated sites. They check pages one at a time. But the SpamBrain risks of pSEO are *between* pages: doorway clusters, near-duplicates, entity-swap templates, thin-content propagation. You can't catch them with per-page rules.\n\npseolint audits the graph: it groups results by template before surfacing them. Run it before you publish, gate it in CI, fix the broken template before SpamBrain does.\n\n### How it compares\n\n|  | pseolint | Screaming Frog | Ahrefs Site Audit | Sitebulb |\n|---|:---:|:---:|:---:|:---:|\n| Unit of analysis | **template cluster** | URL | URL | URL |\n| Near-duplicate / doorway / entity-swap detection | ✅ | partial | - | - |\n| SpamBrain-policy risk verdict | ✅ | - | - | - |\n| AEO / AI-Overview citability checks | ✅ | - | - | - |\n| AI fix → pull request | ✅ | - | - | - |\n| CLI · GitHub Action · MCP server | ✅ | desktop | SaaS | desktop |\n| Open source | ✅ MIT | - | - | - |\n\nThe general-purpose crawlers do plenty pseolint doesn't (JS rendering at scale, backlink data, log-file analysis). pseolint is the specialist for the one thing they weren't built for: **programmatic-SEO compliance at the template level.**\n\n## How pseolint differs\n\n- **Graph-level, not page-level.** Detects near-duplicate clusters, doorway patterns, and entity-swap doorways across thousands of pages. Per-page tools can't see these.\n- **SpamBrain + AI Overview.** 65 rules across 8 categories: SpamBrain-policy mapping (penalty risk) plus `aeo/*` (AI Overview citability: `llms.txt`, AI-crawler access, citable facts, answer-first, summary-bait).\n- **Developer workflow, not SaaS UI.** CLI, GitHub Action, JSON/HTML reports, MCP server, browser extension (SERP competitive recon). Lives in your repo and your PRs.\n- **Actionable, not advisory.** Every finding has a fix, an effort tag (`quick fix` / `moderate` / `structural`), and a Google docs reference.\n- **Safe for hosted use.** SSRF guard (DNS-validated), robots.txt honoured for our own crawler, analytics-blocking in render mode, `AbortSignal` cancellation, `safeMode: \"saas\"` preset for embedding in services.\n- **Calibrated against reputable pSEO.** Engine verdicts are calibrated against a curated corpus of in-production pSEO sites that demonstrably win in search. Doorway-pattern findings cluster (no more per-pair noise); verdicts are reproducible at a fixed `sampleSeed`. Dated snapshot results, the open-source corpus, and the trade-offs we accepted live at [pseolint.dev/methodology](https://pseolint.dev/methodology). Spec: [docs/superpowers/specs/2026-05-03-calibration-against-reputable-pseo.md](./docs/superpowers/specs/2026-05-03-calibration-against-reputable-pseo.md).\n- **Authority-blind by design, with a manual override.** pseolint analyses static content + the link graph it can see. It does NOT measure backlinks, brand mentions, domain age, or any external trust signal: there is no Moz/Ahrefs/Semrush dependency. This means the engine itself is calibrated for the authority tier of the calibration corpus (established brands). It exposes `authorityScore` (0-100, via the `--authority-score` CLI flag, the core API, or the MCP param) so callers can adjust the verdict ladder for their tier: `>= 80` shifts one tier lenient (established brand can absorb shapes a newer site can't); `<= 30` shifts one tier stricter. Raw `risk` number unchanged so CI gates stay stable. Without the flag, treat verdicts as a directional minimum.\n- **Honest about blind spots.** Beyond domain authority, pseolint does not currently detect: image SEO dimensions, schema-content drift (e.g. JSON-LD price ≠ rendered price), outbound-link health, search-intent alignment, parameter-URL crawl-budget waste, and a handful of specialty gaps (mobile-friendliness, cookie-banner detection, AMP/News/Video schema). The complete blind-spot audit lives at [docs/superpowers/specs/2026-05-03-pseolint-blind-spots.md](./docs/superpowers/specs/2026-05-03-pseolint-blind-spots.md): every gap categorized by impact tier with the roadmap fix.\n\nFull version history (calibration rounds, per-rule changes, safety hardening) is in [CHANGELOG.md](./CHANGELOG.md).\n\n## Quick Start\n\n```bash\n# Point it at your local dev server: that's it\nnpx pseolint http://localhost:3000\n```\n\nAutomatically discovers all pages by following internal links. No sitemap, no config, no build step needed.\n\n```bash\n# Save a visual report\nnpx pseolint http://localhost:3000 --format html --output report.html\n\n# Audit a live site (per-template output is the default)\nnpx pseolint https://yoursite.com\n\n# CI gate on build output\nnpx pseolint ./out --ci-threshold concerning --format json\n```\n\n### Per-template output (v0.6 default)\n\n```\nVerdict: CONCERNING\nIntegrity C · Discoverability B · Citation C · Data A\n\nPer-template breakdown (3 templates):\n\n  /listing/:slug  CONCERNING  C\n  10/8201 URLs (0.1%)  uniformity 85%\n  8/10 samples fail `spam/thin-content`\n\n  /category/:slug  READY  A\n  10/312 URLs (3.2%)  uniformity 94%\n\n  /help/:slug  CAUTION  B\n  10/47 URLs (21.3%)  uniformity 78%\n  3/10 samples fail `content/missing-author`\n```\n\n`--format json` includes the `templates` array alongside the existing `findings` list:\n\n```json\n{\n  \"verdict\": \"concerning\",\n  \"risk\": 60,\n  \"templates\": [\n    {\n      \"signature\": \"/listing/:slug\",\n      \"totalUrls\": 8201,\n      \"auditedUrls\": [\"https://example.com/listing/foo\", \"...\"],\n      \"verdict\": \"concerning\",\n      \"risk\": 60,\n      \"variance\": {\n        \"uniformityScore\": 0.85,\n        \"topDriver\": { \"ruleId\": \"spam/thin-content\", \"fireRate\": 0.8 }\n      }\n    },\n    { \"signature\": \"/category/:slug\", \"verdict\": \"ready\", \"risk\": 12 }\n  ],\n  \"findings\": [...]\n}\n```\n\nUse `--legacy-flat` to suppress the template cards and get the v0.5-style flat findings list.\n\n### Partial coverage (`truncated`)\n\nIf the crawl is interrupted (e.g. the backpressure watchdog aborts because the origin is degrading), pseolint still emits whatever it collected, flagged as partial:\n\n```json\n{\n  \"verdict\": \"ready\",\n  \"risk\": 12,\n  \"truncated\": true,\n  \"truncatedReason\": \"Origin degraded mid-crawl (p95 latency exceeded threshold)\",\n  \"pageCount\": 42\n}\n```\n\nWhen `truncated` is `true`, **treat `pageCount`, `risk`, and `verdict` as lower bounds**: a partial pass is not a full pass. The CLI prints a `PARTIAL REPORT` banner and exits non-zero; the GitHub Action warns (and can fail with `fail-on-truncated: true`); the MCP tools and web report surface the same flag. Programmatic consumers should branch on it. The full output contract is published as a JSON Schema (`packages/core/schemas/audit-summary.schema.json`, `$id` carries the `schemaVersion`).\n\n## Audit Modes\n\n| Mode | Command | What you get |\n|------|---------|-------------|\n| **Local dev server** | `npx pseolint http://localhost:3000` | Full rendered pages, HTTP headers, redirect detection, crawl discovery. **Best results.** |\n| **Live site** | `npx pseolint https://yoursite.com` | Same as above against production. Slower (network latency). |\n| **Build directory** | `npx pseolint ./out` | Static HTML files only. No HTTP headers, no redirect detection, no soft-404 detection, no sitemap comparison. Use for CI gates. |\n\n> **Why localhost is recommended:** Build directories contain framework artifacts (Next.js `[slug].html` shells, empty client-rendered pages) that produce false positives. Your dev server renders the actual pages Google will see: with canonicals, meta tags, and full content.\n\n## What It Checks\n\n**65 rules** across **8 categories** (all 8 scored), producing a weighted **SpamBrain Risk Score** (0-100) and an independent **AEO sub-score** for AI Overview citability. Every rule is backed by a primary source (Google Search Central, sitemaps.org, ogp.me, Lighthouse); the checks we deliberately *refuse* to run (folklore the primary sources contradict, like title/description character limits) live in [docs/folklore.md](./docs/folklore.md):\n\n### SpamBrain Risk Detection\n\n| Rule | What It Checks | Severity |\n|------|---------------|----------|\n| `spam/near-duplicate` | SimHash similarity between all page pairs (>85%) | Critical |\n| `spam/entity-swap` | Doorway pages where only a proper noun changes | Critical |\n| `spam/doorway-pattern` | Composite: entity-swap + thin + identical structure + same meta | Critical |\n| `spam/thin-content` | Pages below 300 words (excluding nav/header/footer) | Error |\n| `spam/boilerplate-ratio` | Pages with >70% shared template content | Error |\n| `spam/template-diversity` | Identical DOM structure across all pages | Warning |\n| `spam/publication-velocity` | >100 pages sharing the same publish date | Warning |\n| `spam/template-coverage` | Template dimension coverage (e.g. 87 of 960 possible combinations) | Info |\n\n### Content Quality\n\n| Rule | What It Checks | Severity |\n|------|---------------|----------|\n| `content/unique-value` | Each page must have 100+ words not found on any other page | Error |\n| `content/meta-uniqueness` | Meta descriptions identical after entity masking | Error |\n| `content/title-uniqueness` | Empty/missing title, very short or excessively long title, or two pages sharing the exact title (raw, not entity-masked: catalog templates with per-record entity values pass) | Error / Warning / Info |\n| `content/heading-structure` | No `<h1>`, multiple `<h1>` elements, or long pages (>600 words) with no `<h2>` sub-headings | Error / Warning / Info |\n| `content/image-alt-text` | `<img>` tags missing `alt` attribute (decorative images marked `role=\"presentation\"` / `aria-hidden=\"true\"` / `alt=\"\"` are skipped) | Warning / Info |\n| `content/image-attributes` | `<img>` tags with no width/height (and no inline sizing), which leaves the browser no aspect ratio to reserve and shows up as layout shift; plus pages serving 3+ images where none uses `srcset`/`<picture>` | Warning / Info |\n| `content/missing-author` | No author schema, meta, byline, or rel=\"author\" | Warning |\n| `content/eeat-signals` | Missing E-E-A-T signals (author, dates, sources, about links) | Info |\n| `content/citation-coverage` | Pages making 3+ quantified claims with no authoritative citations | Warning |\n| `content/meta-description-presence` | Missing or empty meta description. Length is deliberately NOT linted: Google documents no character limit (see [docs/folklore.md](./docs/folklore.md)) | Warning |\n\n### Internal Linking\n\n| Rule | What It Checks | Severity |\n|------|---------------|----------|\n| `links/orphan-pages` | Pages with zero inbound internal links | Error |\n| `links/host-section-divergence` | Sub-sections (e.g. `/coupons/`, `/deals/`) that diverge from the rest of the host on ≥2 of: cross-section inbound links, topic vocabulary, template signature, authorship coverage. Targets Google's May 2024 site-reputation-abuse policy. | Warning / Error |\n| `links/dead-ends` | Pages with zero outbound internal links | Warning |\n| `links/cluster-connectivity` | Isolated page clusters with no cross-linking | Warning |\n| `links/unreachable-from-root` | Pages with no path from the start URL (graph-disconnected from the entry point) | Warning |\n| `links/link-depth` | Pages requiring >3 clicks from root | Info |\n| `links/crawlable-anchors` | Links Google cannot follow: `<a>` without `href`, `javascript:` hrefs, onclick/router-attribute pseudo-links. Escalates to Error when a page's navigation is effectively invisible to crawlers | Warning / Error |\n| `links/generic-anchor-text` | ≥50% of a page's internal links anchored on \"click here\" / \"read more\" / empty text, wasting the anchor signal Google (and AI answer engines) use to label the target | Info |\n\n### Technical SEO\n\n| Rule | What It Checks | Severity |\n|------|---------------|----------|\n| `tech/canonical-consistency` | Missing, invalid, or conflicting canonical URLs (HTML + HTTP header) | Error |\n| `tech/sitemap-completeness` | Pages missing from sitemap, phantom 404s, redirecting sitemap URLs | Error |\n| `tech/csr-bailout` | Render-diff: substantive content / interactivity that appears only after client-side JS: invisible to crawlers and the first indexing pass (needs `--render`) | Warning |\n| `tech/core-web-vitals` | Core Web Vitals in Google's \"poor\" tier. Default: lab LCP/CLS from a headless-Chromium render (needs `--render`). With a free CrUX API key (`--crux-api-key`), uses real-user field p75 for LCP/CLS **and INP**: the numbers Google ranks on | Warning |\n| `tech/soft-404` | HTTP 200 pages that look like error pages: plus a synthetic-URL probe that fetches one nonexistent URL per template cluster (a 200 means the directory will index unbounded junk; needs `--render`) | Error |\n| `tech/robots-compliance` | Sitemap URLs blocked by `robots.txt` (Disallow patterns matching listed pages) | Error |\n| `tech/robots-noindex-conflict` | Noindexed pages (meta or X-Robots-Tag) with inbound links | Warning |\n| `tech/canonical-noindex-conflict` | Noindex + canonical pointing elsewhere | Warning |\n| `tech/redirect-chain` | Redirect chains longer than 2 hops | Warning |\n| `tech/hreflang-consistency` | Hreflang reciprocity (A->B requires B->A) | Warning |\n| `tech/og-completeness` | Missing `og:title`, `og:description`, or `og:image`: affects social-share previews and AI Overview fallback summaries | Warning |\n| `tech/robots-sitemap-presence` | Missing or unreachable `/robots.txt` or `/sitemap.xml` at the origin | Warning |\n| `tech/language-mismatch` | Declared language (html lang / self-referencing hreflang) vs the Unicode script of the actual text, e.g. `lang=\"ja\"` on a Cyrillic page. Google indexes by DETECTED language, so mismatched declarations silently break all targeting | Error / Warning / Info |\n| `tech/hreflang-validity` | Invalid hreflang codes (`en_US`, `jp`, `en-UK`); Google silently ignores the whole annotation | Warning |\n| `tech/html-size` | HTML approaching Googlebot's 2 MB per-file crawl cutoff (uncompressed; content/links/JSON-LD past it are invisible). Per-file, not total page weight (see [docs/folklore.md](./docs/folklore.md)) | Error / Warning |\n| `tech/resource-weight` | Subresource bytes from the `--render` pass: any single script/style/image at or past the same 2 MB per-file cutoff (Error), within 25% of it (Warning), plus a total-page-weight breakdown by kind (Info, explicitly not a crawl limit) | Error / Warning / Info |\n| `tech/meta-robots-conflict` | Contradictory robots directives across meta robots / meta googlebot / X-Robots-Tag. Google applies the MOST restrictive, so an accidental `noindex` silently wins | Error / Warning |\n| `tech/snippet-suppression` | `nosnippet` / `max-snippet:0`, which kills SERP snippets and AI Overview / answer-engine citability | Warning / Info |\n| `tech/viewport-meta` | Missing `<meta name=\"viewport\">`; Google indexes mobile-first | Warning |\n| `tech/sitemap-hygiene` | Cross-host sitemap URLs (dropped per sitemaps.org), future / unparseable / mass-identical `lastmod` values (Google ignores unreliable lastmod) | Error / Warning |\n| `tech/robots-txt-limits` | robots.txt over Google's 500 KiB parse limit, or unsupported directives (`noindex:` in robots.txt has been ignored since 2019, so pages are NOT excluded) | Warning / Info |\n\n### Data Consistency\n\n| Rule | What It Checks | Severity |\n|------|---------------|----------|\n| `data/missing-binding` | When `--data-source` is set, flags fields from the source record that don't appear on the matching page (e.g. FAQ items, regulation clauses listed in the source JSON but missing from rendered HTML) | Warning |\n| `data/identical-across-pages` | Source-data fields that differ in the JSON but render identically across pages (suggests a missing binding loop or a hardcoded template value) | Warning |\n\n### Structured Data\n\n| Rule | What It Checks | Severity |\n|------|---------------|----------|\n| `schema/json-ld-valid` | Malformed JSON-LD, missing @context or @type | Error |\n| `schema/required-fields` | Article/Product/FAQ missing required fields | Warning |\n| `schema/consistency` | Mixed schema types across template pages | Info |\n\n### Cannibalization\n\n| Rule | What It Checks | Severity |\n|------|---------------|----------|\n| `cannibal/url-pattern` | URL structures with same tokens in different order | Info |\n\n> `cannibal/title-overlap` and `cannibal/keyword-collision` were dropped in v0.4 due to high false-positive rates on legitimately similar pages (e.g. localized variants, paginated archives). See the [v0.4 redesign spec §4.3](./docs/superpowers/specs/2026-04-29-pseolint-v0.4-engine-redesign.md).\n\n### AEO: AI Overview Readiness (v0.3.x)\n\n| Rule | What It Checks | Severity |\n|------|---------------|----------|\n| `aeo/llms-txt` | `/llms.txt` missing or malformed at the origin | Warning |\n| `aeo/crawler-access` | `robots.txt` blocks `GPTBot` / `ClaudeBot` / `PerplexityBot` / `Bytespider` / `Google-Extended` / `CCBot` / `Applebot-Extended` / `ChatGPT-User` | Warning / Error |\n| `aeo/freshness-signals` | No `dateModified` / modification meta / visible \"Last updated\" | Warning |\n| `aeo/faq-coverage` | FAQ-style content (question-phrased H2s) without `FAQPage` / `HowTo` JSON-LD | Info |\n| `aeo/answer-first` | First paragraph after H1 is boilerplate or lacks facts / named entities | Error |\n| `aeo/citable-facts` | <3 entity-specific citable facts per page after template-fact filtering | Error |\n| `aeo/content-modularity` | Sections that cross-reference each other or use vague headings: not independently extractable | Warning |\n| `aeo/summary-bait` | Composite: strong opener + no interactive value + facts packed in opener → guaranteed zero-click loss | Error |\n\n## Live URL Scanning\n\nWhen you point pseolint at a URL, it captures what Google sees:\n\n- **HTTP metadata**: status codes, redirect chains, X-Robots-Tag, Link headers\n- **Crawl discovery**: follows internal links from the start page to find all crawlable pages\n- **Sitemap comparison**: if a sitemap exists, compares it against crawl-discovered pages\n\n```bash\n# Just give it your homepage: it discovers everything\nnpx pseolint https://paperforge.dev\n```\n\n## Page Groups\n\nDifferent page types need different standards. Configure groups in `pseolint.config.ts`:\n\n```typescript\nexport default {\n  pageGroups: {\n    pseo: {\n      match: '/templates/**',\n      rules: ['spam/*', 'content/*', 'links/*', 'cannibal/*', 'tech/*', 'schema/*'],\n      overrides: {\n        'spam/thin-content': { thinContentMinWords: 500 },\n      }\n    },\n    listing: {\n      match: ['/documents', '/templates'],\n      rules: ['tech/*'],\n    },\n    marketing: {\n      match: ['/', '/about', '/pricing'],\n      rules: ['tech/*'],\n    },\n    utility: {\n      match: ['**/404*', '**/500*'],\n      rules: [],  // skip entirely\n    }\n  }\n};\n```\n\nEach group gets its own score. Unmatched pages get all rules.\n\n## SpamBrain Risk Score\n\nThe risk score (0–100) aggregates rule penalties into **4 super-categories** (**Integrity** (spam + content + cannibal), **Discoverability** (links + tech), **Citation** (aeo + schema), and **Data**) with site-type-aware weights, so a programmatic directory and a docs site are each scored against the rule weighting that matches their archetype. Since v0.6, scoring runs **per template** and rolls up to a site verdict: the worst-scoring template that covers ≥5% of the audited URLs.\n\nThe score maps to a 4-rung **verdict ladder**, and CI gates on the verdict (`--ci-threshold`, default `concerning`), not a raw numeric band:\n\n| Verdict | Meaning | CI exit (verdict ≥ threshold) |\n|---------|---------|-------------------------------|\n| `ready` | no material risk | 0 |\n| `caution` | minor issues | 0 |\n| `concerning` | likely penalty-pattern exposure | 1 |\n| `critical` | strong penalty-pattern exposure | 1 |\n\nSee [pseolint.dev/methodology](https://pseolint.dev/methodology) for the calibrated weights and verdict thresholds.\n\n## Actionable Output\n\nFindings are automatically enriched before display:\n\n- **Pairwise clustering**: Thousands of near-duplicate pair comparisons collapse into a handful of cluster findings: \"48 pages form a near-duplicate cluster (86–94% similar).\"\n- **Content breakdown**: Each cluster shows what's shared vs. unique: \"Shared: description of property (31w), buyer acknowledges (35w). Unique: 3324w of 8140w.\"\n- **Effort tags**: Every finding is tagged `quick fix`, `moderate`, or `structural` so you know where to start.\n- **Template detection**: When the tool detects template-generated content, fix suggestions speak to template authors: \"Add conditional content sections per entity.\"\n\n## CLI Options\n\n```\nUsage: pseolint [options] [command] [source]\n\nArguments:\n  source                         URL or directory path to audit\n\nOutput\n  -f, --format <type>            Output format: console | json | markdown | html (default: console)\n  --ci-threshold <severity>      Min verdict that fails CI: ready|caution|concerning|critical (default: concerning)\n  -t, --threshold <n>            [deprecated] Numeric risk threshold; use --ci-threshold instead\n  -o, --output <file>            Write report to file instead of stdout\n  --no-color                     Disable colored output\n\nCrawl / fetch\n  --concurrency <n>              Max parallel HTTP fetches (default: 5)\n  --timeout <ms>                 Per-request timeout in ms (default: 30000)\n  --no-crawl                     Disable crawl-based page discovery for URL sources\n  --ignore <patterns>            Comma-separated glob patterns to exclude\n  --render                       Render pages in a browser before auditing\n  --browser-ws <url>             CDP WebSocket endpoint for browser rendering\n\nSampling\n  --sample-size <n>              Audit N pages (default: 0 = all)\n  --strategy <random|stratified> Sampling strategy (default: stratified)\n  --max-per-template <n>         Cap samples per URL template cluster (default: 0)\n\nTemplate output (v0.6)\n  --per-template                 Render per-template cards above the findings list (default: ON)\n  --template <signature>         Filter output to a single template, e.g. /listing/:slug\n  --legacy-flat                  Suppress template cards; print the v0.5-style flat findings list\n\nCache & monitoring\n  --cache [dir]                  Enable HTTP cache (default: .pseolint/cache)\n  --cache-ttl <duration>         TTL for entries without validators, e.g. 7d, 1h, 30m (default: 7d)\n  --state [path]                 Enable state persistence (default: .pseolint/state.json)\n  --mode <monitoring|fresh>      v0.5+ change-driven monitoring mode. Auto-monitoring is the\n                                 default when prior state exists. Use 'fresh' to force a full\n                                 re-audit even with prior state.\n  --age-floor-days <n>           v0.5+ minimum days since a URL's last fetch before monitoring\n                                 forces a re-fetch regardless of other signals (default: 7)\n  --since                        v0.5+ alias for --mode=monitoring (kept for back-compat)\n  --exit-on-regression           Exit non-zero when new rule IDs fire vs prior --state\n\nData\n  --data-source <file>           JSON file with source data for content-verification rules\n\nAI triage (opt-in)\n  --ai                           Enable AI triage of findings\n  --ai-provider <id>             anthropic | openai | google | mistral | groq | xai | cohere | ollama\n  --ai-model <name>              Model name (overrides provider default)\n  --ai-endpoint <url>            AI endpoint (Ollama only; default: http://localhost:11434)\n  --ai-max-tokens <n>            Input token cap per triage call (default: 60000)\n  --ai-max-cost <usd>            Refuse a triage call whose pre-flight cost exceeds this USD\n  --ai-daily-budget <usd>        Refuse triage when today's total spend would exceed USD (requires --telemetry)\n  --ai-cache-ttl <duration>      Triage cache TTL, e.g. 30d, 12h, 60s (default: 30d)\n  --no-ai-cache                  Bypass AI triage cache for this run\n  --no-ai-suggest                Suppress AI discovery hint in non-AI runs\n\nTelemetry (local, offline)\n  --telemetry                    Enable local telemetry write (.pseolint/telemetry.jsonl)\n  --telemetry-path <file>        Override telemetry JSONL path\n  --no-telemetry-prompt          Suppress the y/n/skip triage feedback prompt\n  --triage-feedback <rating>     Non-interactive feedback: helpful | unhelpful | y | n\n\nMCP\n  --mcp                          Start as an MCP server (for AI coding assistants)\n\nCommands:\n  stats                          Show aggregate telemetry stats from .pseolint/telemetry.jsonl\n  stats-export <outPath>         Copy telemetry JSONL to <outPath> for manual review/sharing\n```\n\n### Caching & change-driven monitoring (v0.5)\n\n```bash\n# First run: populates .pseolint/cache and .pseolint/state.json with full baseline\nnpx pseolint https://yoursite.com --cache --state\n\n# Subsequent runs auto-enter monitoring mode. The decision matrix decides which\n# URLs to fetch BEFORE the network round-trip:\n#   - new URL                         → fetch (reason: new)\n#   - prior fetch ≥ 7 days old        → fetch (reason: age)\n#   - ruleset version bumped          → fetch (reason: ruleset)\n#   - prior warning/error finding     → fetch (reason: recheck): info findings carry forward\n#   - sitemap <lastmod> newer         → fetch (reason: lastmod)\n#   - none of the above + lastmod present → SKIP (carry findings forward)\nnpx pseolint https://yoursite.com --cache --state\n\n# Force a full re-audit even with prior state\nnpx pseolint https://yoursite.com --cache --state --mode=fresh\n\n# Lower the age-floor for tighter monitoring (default: 7 days)\nnpx pseolint https://yoursite.com --cache --state --age-floor-days=3\n\n# CI gate that fails when a *new* rule ID starts firing on actually-fetched URLs\nnpx pseolint https://yoursite.com --cache --state --exit-on-regression\n```\n\nSites whose sitemaps emit `<lastmod>` (Next.js, Yoast/WordPress, Astro) get the\nbiggest savings: typically ~95% fewer fetches on steady-state monitoring runs.\nSites without `<lastmod>` hit `no-signal` and refetch every URL; bandwidth is\nstill saved via cache.ts conditional GETs but round-trips aren't skipped (a\nHEAD-fallback path is on the roadmap).\n\nEnd-of-run summary line:\n```\nMonitoring: 47/4012 URLs re-scraped (recheck=23, lastmod=12, age=8, new=4), 3965 carried forward.\n```\n\n### AI triage\n\nTurns hundreds of findings into a handful of ranked root causes. Opt-in, bring-your-own API key, with cost guardrails:\n\n```bash\n# Auto-detect provider from env (ANTHROPIC_API_KEY, OPENAI_API_KEY, etc.)\nnpx pseolint https://yoursite.com --ai\n\n# Pin provider + model, cap spend\nnpx pseolint https://yoursite.com --ai \\\n  --ai-provider anthropic \\\n  --ai-model claude-haiku-4-5 \\\n  --ai-max-cost 0.50\n\n# Local-only (Ollama, no network cost)\nnpx pseolint https://yoursite.com --ai --ai-provider ollama --ai-model qwen2.5:7b\n\n# Enforce a daily spend ceiling across runs (requires telemetry)\nnpx pseolint https://yoursite.com --ai --telemetry --ai-daily-budget 5.00\n```\n\nEvery call prints a pre-flight cost estimate before hitting the provider. Cache hits don't count against the daily budget.\n\n### Local telemetry & stats\n\nTelemetry is **local JSONL only**: zero network, counts + spend + feedback ratings. Off by default.\n\n```bash\nnpx pseolint https://yoursite.com --ai --telemetry\nnpx pseolint stats              # show your success rate, spend, feedback ratio\nnpx pseolint stats-export out.jsonl  # copy log for manual inspection\n```\n\n## Browser Rendering\n\nFor client-rendered sites (React SPAs, Next.js app router), use `--render` to capture the fully rendered DOM:\n\n```bash\n# With a remote CDP endpoint (Browserless, etc.)\nPSEOLINT_BROWSER_WS=wss://your-browser:3000 npx pseolint https://yoursite.com --render\n\n# With local Playwright\nnpm install playwright-core\nnpx playwright install chromium\nnpx pseolint https://yoursite.com --render\n```\n\nWorks with any CDP-compatible browser. Remote endpoints must use `wss://`.\n\n## Core Web Vitals\n\nTwo sources, both opt-in:\n\n```bash\n# Lab: measure LCP + CLS from a headless-Chromium render. Zero external calls.\nnpx pseolint https://yoursite.com --render\n\n# Field: real-user p75 LCP/CLS/INP from the Chrome UX Report (the numbers Google\n# ranks on, and the only source of INP). Free key: https://developer.chrome.com/docs/crux/api\nCRUX_API_KEY=... npx pseolint https://yoursite.com          # or --crux-api-key <key>\n\n# Query the mobile field data specifically (Google indexes mobile-first)\nCRUX_API_KEY=... npx pseolint https://yoursite.com --crux-form-factor phone\n```\n\nSelection is **per-metric**: when a CrUX key is set, `tech/core-web-vitals` uses field\ndata for each of LCP/CLS/INP and falls back to the lab render for any metric CrUX\nlacks, so enabling field data never drops a signal the lab render already had.\n\nCrUX only covers URLs/origins with enough real traffic, so low-traffic pSEO pages get\ntheir **origin-level** field vitals as a fallback. A site-wide origin reading collapses\ninto **one** finding (not one per page). Per-URL lookups are pooled and capped at 150\n(`--crux-max-lookups <n>`, or `0` for unlimited); if the cap forces origin-level\nfallback, or CrUX rate-limits (429) / rejects the key (401/403), pseolint says so rather\nthan silently reporting \"no data\". The CrUX endpoint is a fixed Google host, so no\nexternal-authority dependency on your own content, consistent with pseolint's\noffline-runnable design.\n\n## GitHub Action\n\n```yaml\nname: pSEO Lint\non: [pull_request]\n\njobs:\n  audit:\n    runs-on: ubuntu-latest\n    steps:\n      - uses: actions/checkout@v4\n      - uses: actions/setup-node@v4\n        with: { node-version: '20' }\n      - run: npm run build\n      - uses: ouranos-labs/pseolint@v0.8.0\n        with:\n          source: ./out\n          threshold: 40\n```\n\nPosts a score summary as a PR comment and fails the check if score exceeds the threshold.\n\nPin to a release tag (`@v0.8.0`) or, if you prefer a moving ref that tracks the\nlatest built bundle, `ouranos-labs/pseolint/packages/action@action-v1`. The tag\nform is the one listed on the GitHub Actions Marketplace: Marketplace only\nindexes an `action.yml` at the repository root, so `/packages/action@…` resolves\nthe same action but is invisible to search there.\n\n| Input | Required | Default | |\n| --- | --- | --- | --- |\n| `source` | yes | - | Build output directory or sitemap URL |\n| `threshold` | no | `40` | Risk score at or above which the check fails |\n| `token` | no | `github.token` | Token used to post the PR comment |\n| `fail-on-truncated` | no | `false` | Fail when coverage was partial, even if risk is under threshold |\n\nOutputs: `risk`, `score` (alias), `verdict`, `pageCount`, `truncated`,\n`truncated-reason`. When `truncated` is `true`, treat the numbers as lower\nbounds: see [Partial coverage](#partial-coverage-truncated).\n\n## Fix rail: from audit to pull request\n\nThe AI orchestrator produces a **fix manifest** (validated patches). `pseolint apply` writes the deterministic ones (meta titles, H1s, `robots.txt`, `sitemap.xml`) straight into your source tree; generative or unmatched patches are demoted to a checklist for a human. `--pr` takes the next step: commit those edits to a tool-owned branch and open a PR.\n\n```bash\n# 1. Audit → manifest\npseolint orchestrate https://example.com --max-cost 3 --manifest-out manifest.json\n\n# 2. Apply deterministic edits into your working tree (review the diff, commit yourself)\npseolint apply manifest.json\n\n# 3. …or apply + commit + open a GitHub PR in one step\npseolint apply manifest.json --pr --token \"$GITHUB_TOKEN\"\n```\n\n### Mapping (`.pseolint/templates.json`)\n\nAudited routes don't know your source layout, so you map them once (route pattern → source file). Domain-level patches use the special `robots.txt` / `sitemap.xml` keys:\n\n```json\n{\n  \"/listing/:slug\": \"app/listing/[slug]/page.tsx\",\n  \"/category/:slug\": \"app/category/[slug]/page.tsx\",\n  \"robots.txt\": \"public/robots.txt\",\n  \"sitemap.xml\": \"app/sitemap.ts\"\n}\n```\n\nRoute keys accept `:seg` / `[seg]` / `*` wildcards. A patch with no matching entry (or a literal that can't be found in an interpolated template like `Best in ${city}`) lands in the checklist (or the PR body) instead of silently corrupting source.\n\n### In CI\n\n`apply --pr` uses `git` + one GitHub API call; no extra dependency. Give the workflow write permissions and let `actions/checkout` configure the push token:\n\n```yaml\nname: pSEO fix PR\non: { workflow_dispatch: {} }\n\njobs:\n  fix:\n    runs-on: ubuntu-latest\n    permissions: { contents: write, pull-requests: write }\n    steps:\n      - uses: actions/checkout@v4\n      - uses: actions/setup-node@v4\n        with: { node-version: '20' }\n      - run: npm run build\n      - run: npx pseolint orchestrate http://localhost:3000 --max-cost 3 --manifest-out manifest.json\n        env: { ANTHROPIC_API_KEY: '${{ secrets.ANTHROPIC_API_KEY }}' }\n      - run: npx pseolint apply manifest.json --pr\n        env: { GITHUB_TOKEN: '${{ github.token }}' }\n```\n\nRe-running updates the same `pseolint/fix-<domain>` branch (force-with-lease, tool-owned branch only); it never spams new PRs. It no-ops cleanly when there's nothing deterministic to apply.\n\n## Output Formats\n\n```bash\nnpx pseolint https://yoursite.com                  # Colored terminal (default)\nnpx pseolint https://yoursite.com --format json    # CI-friendly JSON\nnpx pseolint https://yoursite.com --format markdown # PR comments / docs\nnpx pseolint https://yoursite.com --format html    # Self-contained visual report\n```\n\n## Monorepo\n\n| Package | npm | Version | License |\n|---------|-----|---------|---------|\n| `packages/core` | [`@pseolint/core`](packages/core/README.md) | 0.7.5 | MIT |\n| `packages/cli` | [`pseolint`](packages/cli/README.md) | 0.7.3 | MIT |\n| `packages/mcp` | [`@pseolint/mcp`](packages/mcp/README.md) | 0.7.4 | MIT |\n| `packages/action` | GitHub Action (`ouranos-labs/pseolint@v0.8.0`, Marketplace) | - | MIT |\n| `apps/web` | pseolint.dev | - | AGPL-3.0 |\n\n## Development\n\n```bash\nbun install\nbun run build\nbun run test     # 1,203 tests across 126 files (core)\n```\n\n## Roadmap\n\n- **AI-inferred template mapping**: today `apply --pr` needs a hand-authored `.pseolint/templates.json`; infer route→source automatically.\n- **Closing blind spots**: schema-content drift, outbound-link health, search-intent alignment. Every gap is tracked by impact tier in the [blind-spot audit](./docs/superpowers/specs/2026-05-03-pseolint-blind-spots.md). (Core Web Vitals landed: lab LCP/CLS via `--render`, real-user field p75 + INP via `--crux-api-key`.)\n- **Web \"Open PR\" button**: the fix rail runs from the CLI/Action today; a hosted one-click flow is deferred until the GitHub-App auth is justified.\n\nFound a false positive or a missing check? [Open an issue](https://github.com/ouranos-labs/pseolint/issues). Corpus-backed bug reports move the calibration.\n\n## Contributing\n\nIssues and PRs welcome. See [CONTRIBUTING.md](./CONTRIBUTING.md) for the dev loop, and [`skills/`](skills/README.md) if you want to teach an agent to design pass-first pages. If pseolint saved you a SpamBrain headache, a ⭐ helps others find it.\n\n## License\n\nMIT (packages) / AGPL-3.0 (apps/web)\n",
  "bytes": 38809,
  "sha": "b7d080d9d1fc8ab094fb3739288ae3cf3f83730cd1f6075e6a02b25024f9fab9",
  "repo_slug": "ouranos-labs/pseolint",
  "fonte": "repo",
  "truncated": false,
  "api": "https://agentalog.com/api/listings/mcp_io_github_ouranos_labs_pseolint_859754d1/readme"
}