{
  "markdown": "# RigorPilot Skills\n\nRun research repositories from their README, with bounded execution and auditable evidence.\nRigorPilot adds section-level results without rewriting the original README.\nTrusted reproduction is the default; candidate exploration requires explicit authorization.\n\n[English](README.md) · [简体中文](README.zh-CN.md)\n\n[![Skillselion Top 100](https://skillselion.com/badge/skills/lllllllama/rigorpilot-skills/paper-context-resolver.svg?award=1)](https://skillselion.com/skills/lllllllama/rigorpilot-skills/paper-context-resolver)\n\n<p>\n  <a href=\"https://github.com/lllllllama/RigorPilot-Skills/actions/workflows/validate.yml\"><img alt=\"CI\" src=\"https://github.com/lllllllama/RigorPilot-Skills/actions/workflows/validate.yml/badge.svg\"></a>\n  <a href=\"https://skillselion.com/skills/lllllllama/rigorpilot-skills/ai-research-reproduction\"><img alt=\"Listed on Skillselion\" src=\"https://skillselion.com/badge/skills/lllllllama/rigorpilot-skills/ai-research-reproduction.svg\"></a>\n  <a href=\"https://skills.sh/lllllllama/rigorpilot-skills\"><img alt=\"skills.sh installs\" src=\"https://skills.sh/b/lllllllama/rigorpilot-skills\"></a>\n  <a href=\"https://github.com/lllllllama/RigorPilot-Skills/stargazers\"><img alt=\"GitHub stars\" src=\"https://img.shields.io/github/stars/lllllllama/RigorPilot-Skills?style=flat-square\"></a>\n  <a href=\"LICENSE\"><img alt=\"MIT License\" src=\"https://img.shields.io/badge/license-MIT-yellow?style=flat-square\"></a>\n  <a href=\"https://agentskills.io\"><img alt=\"Agent Skills standard\" src=\"https://img.shields.io/badge/Agent%20Skills-open%20standard-1f6feb?style=flat-square\"></a>\n  <img alt=\"platforms\" src=\"https://img.shields.io/badge/Windows%20%7C%20Linux-supported-6f42c1?style=flat-square\">\n  <img alt=\"local regression\" src=\"https://img.shields.io/badge/local%20regression-69%2F69%20passed-8250df?style=flat-square\">\n  <a href=\"benchmark_outputs/external_suite_latest.json\"><img alt=\"historical external protocols\" src=\"https://img.shields.io/badge/historical%20protocols-4%2F4%20passed-238636?style=flat-square\"></a>\n</p>\n\n<p align=\"center\">\n  <a href=\"#examples\"><strong>Real examples</strong></a> ·\n  <a href=\"#quick-start\"><strong>Install & use</strong></a> ·\n  <a href=\"#skills\"><strong>Skill index</strong></a> ·\n  <a href=\"#validation\"><strong>Validation</strong></a> ·\n  <a href=\"docs/ENGINEERING_ROADMAP.md\"><strong>Engineering roadmap</strong></a>\n</p>\n\n<a id=\"examples\"></a>\n<a id=\"evidence\"></a>\n\n## 📄 Real repositories, inspectable results\n\nOriginal commands, prose, badges, images, videos and HTML stay in the source file.\nRigorPilot splits that file into sections and inserts one evidence-linked card per section.\nRemoving its insertion blocks restores the retained original README byte for byte.\n\nEach card below opens a full annotated README **beside the original README in a\nretained repository checkout**. Supporting repository files are kept so relative\nlinks and media retain their original context.\n\n🟢 selected checks passed · 🔵 not executed · ⚪ read only · 🟡 partial · 🔴 blocked · 🟣 decision needed.\nGreen does not automatically mean paper-result reproduction; blue is not an execution failure.\n\n<table>\n  <tr>\n    <td align=\"center\" width=\"50%\">\n      <a href=\"benchmark_outputs/showcases/micrograd/repo/RIGORPILOT_README.md\"><img src=\"assets/showcase/external-micrograd.png\" width=\"100%\" alt=\"micrograd: recorded pytest execution and section-level evidence\"/></a><br/>\n      <b>micrograd · correctness checks</b><br/>\n      <sub>🟢 2 tests passed in 7.62 s<br/>8 headings = 8 annotations · original bytes preserved</sub><br/>\n      <a href=\"benchmark_outputs/showcases/micrograd/repo/RIGORPILOT_README.md\">Open full RigorPilot README →</a>\n    </td>\n    <td align=\"center\" width=\"50%\">\n      <a href=\"benchmark_outputs/showcases/mingpt/repo/RIGORPILOT_README.md\"><img src=\"assets/showcase/external-mingpt.png\" width=\"100%\" alt=\"minGPT: target selection only, with no model download or execution\"/></a><br/>\n      <b>minGPT · selection boundary</b><br/>\n      <sub>🔵 Test selected, not executed · no model download<br/>11 headings = 11 annotations · original bytes preserved</sub><br/>\n      <a href=\"benchmark_outputs/showcases/mingpt/repo/RIGORPILOT_README.md\">Open full RigorPilot README →</a>\n    </td>\n  </tr>\n  <tr>\n    <td align=\"center\" width=\"50%\">\n      <a href=\"benchmark_outputs/showcases/pytorch-mnist/repo/mnist/RIGORPILOT_README.md\"><img src=\"assets/showcase/external-pytorch-mnist.png\" width=\"100%\" alt=\"PyTorch MNIST: partial bounded training and captured loss\"/></a><br/>\n      <b>PyTorch MNIST · bounded startup</b><br/>\n      <sub>🟡 Partial training · observed loss 0.038893<br/>1 heading = 1 annotation · original bytes preserved</sub><br/>\n      <a href=\"benchmark_outputs/showcases/pytorch-mnist/repo/mnist/RIGORPILOT_README.md\">Open full RigorPilot README →</a>\n    </td>\n    <td align=\"center\" width=\"50%\">\n      <a href=\"benchmark_outputs/showcases/nanogpt-shakespeare/repo/RIGORPILOT_README.md\"><img src=\"assets/showcase/external-nanogpt.png\" width=\"100%\" alt=\"nanoGPT Shakespeare: partial CPU training and captured train and validation losses\"/></a><br/>\n      <b>nanoGPT Shakespeare · bounded training</b><br/>\n      <sub>🟡 Partial · train loss 4.1676 · validation loss 4.1649<br/>11 headings = 11 annotations · original bytes preserved</sub><br/>\n      <a href=\"benchmark_outputs/showcases/nanogpt-shakespeare/repo/RIGORPILOT_README.md\">Open full RigorPilot README →</a>\n    </td>\n  </tr>\n</table>\n\n[All four cases and upstream links](benchmark_outputs/EXTERNAL_REPRODUCTIONS.md) ·\n[Recorded suite](benchmark_outputs/external_suite_latest.json) ·\n[Case definitions](benchmarks/external_cases.json) · [Methodology](benchmarks/README.md)\n\nThese are historical, commit-pinned deterministic runs: **4/4 case protocols**\npassed in `251.0 s`, with a peak workspace of `98.67 MiB` and `0` model API calls.\nThe zero-API count applies only to that suite. Selection-only and partial cases\nare not completed evaluations, converged training or reproduced paper scores.\n\nNew: [installed-skill micrograd trial](docs/FIRST_USE_ACCEPTANCE.md), with before/after command reports, a retained failed attempt and independent checks—not a model-quality comparison.\n\n<a id=\"quick-start\"></a>\n\n## 🚀 Install and use\n\nThe installer needs Node.js/npm; check its Node version requirement if it reports `EBADENGINE`.\n\nInstall all skills:\n\n```bash\nnpx skills add lllllllama/rigorpilot-skills --all\n```\n\nOr install only the self-contained reproduction skill:\n\n```bash\nnpx skills add lllllllama/rigorpilot-skills --skill ai-research-reproduction\n```\n\nOpen the target repository in a Skills-capable agent, then ask:\n\n> Use ai-research-reproduction: run the smallest README-documented evaluation, preserve the source and write evidence to repro_outputs/, plus an annotated copy beside the original README. Ask before large downloads or long training.\n\nThe main skill works alone; choose **all skills** for companion and leaf entrypoints.\nYour existing agent loads the skill. The standalone model runner is optional.\n[Client compatibility](references/client-compatibility-policy.md)\n\nStart with the `RIGORPILOT_README.md` reported in `source_adjacent_readme.path`,\nthen follow its command and log links. If a conflicting file blocks the extra\ncopy, that file stays intact; inspect `repro_outputs/SUMMARY.md` for the outcome\nand next action.\n\n## What it does—and does not do\n\nREADME → documented target → reviewed setup → bounded execution → verification → evidence.\n\n- Preserves source meaning; records assumptions, deviations, failures and blockers.\n- Records process state, logs and attempt lineage; supports explicit cancellation,\n  recovery and retry through the persistent runtime.\n- Separates trusted reproduction from explicitly authorized, candidate-only exploration.\n- Checks execution criteria independently of the model's completion claim.\n\nThis is **local execution, not an OS sandbox**. Approved commands can access the\nhost and network; use trusted repositories. Resource admission and between-action\nbudget checks are not hard OS quotas or subscription-balance monitoring.\n\nThe optional model loop currently supports Anthropic Messages and reviewed command\nIDs, not unrestricted source repair. **This standalone runner has no successful\nlive-model acceptance recorded yet**: three provider attempts returned HTTP 502. Other model profiles\nare metadata, not proof of working transports or equivalent model performance.\n[Runner and recovery contract](skills/ai-research-reproduction/references/agent-runner.md) ·\n[Implementation evidence and limits](docs/P0_P1_DELIVERY.md)\n\n<a id=\"skills\"></a>\n\n## 🎯 Skill index\n\n| Task | Skill |\n|---|---|\n| Reproduce from README commands | [`ai-research-reproduction`](skills/ai-research-reproduction/SKILL.md) |\n| Read-only repository analysis | [`analyze-project`](skills/analyze-project/SKILL.md) |\n| Prepare environment, data and weights | [`env-and-assets-bootstrap`](skills/env-and-assets-bootstrap/SKILL.md) |\n| Run documented inference or evaluation | [`minimal-run-and-audit`](skills/minimal-run-and-audit/SKILL.md) |\n| Start or verify training conservatively | [`run-train`](skills/run-train/SKILL.md) |\n| Diagnose before proposing a patch | [`safe-debug`](skills/safe-debug/SKILL.md) |\n| Coordinate authorized candidate exploration | [`ai-research-explore`](skills/ai-research-explore/SKILL.md) |\n| Implement a candidate change on an isolated branch | [`explore-code`](skills/explore-code/SKILL.md) |\n| Execute a bounded candidate experiment | [`explore-run`](skills/explore-run/SKILL.md) |\n\nTwo helpers support orchestration: `repo-intake-and-plan` and `paper-context-resolver`.\nExploration requires a durable `current_research` anchor and a frozen comparison\ncontract. Candidate results never become trusted baseline results by declaration.\n[Routing](references/routing-policy.md) · [Research loop](references/research-thinking-loop.md) ·\n[Campaign inputs](skills/ai-research-explore/references/research-campaign-spec.md)\n\n## 📦 Evidence bundle\n\n| Artifact | What to inspect |\n|---|---|\n| `repro_outputs/ANNOTATED_README.md` | Original README with inserted section verdicts |\n| `SUMMARY.md`, `COMMANDS.md`, `LOG.md`, `status.json` | Outcome, exact commands, observations and machine-readable status |\n| `PATCHES.md`, `SCIENTIFIC_CHANGELOG.md`, `COMPARABILITY_REPORT.md` | Changes, scientific meaning and comparison boundaries |\n| `_runtime/<run_id>/` | Process state, events, resource samples and stdout/stderr |\n| `agent_state.json`, `trajectory.jsonl` | Optional model runner's checkpoints, tool calls and reported usage |\n\n🟢 success · 🔵 not executed · ⚪ read only · 🟡 partial · 🔴 blocked · 🟣 decision required\n\nStandard evidence stays under `repro_outputs/`. Both main runners accept\n`--source-adjacent-readme` to also write `RIGORPILOT_README.md` beside the original,\npreserving the context of its relative media/file links. Only inserted evidence\nlinks are rebased. The same output directory may refresh its unchanged owned\ncopy, never an unrelated or manually edited file. Retain supporting repository\nfiles and the evidence directory's `readme_delivery.json`.\n[Output contract](references/output-contract.md) · [Rigor principles](references/research-rigor-principles.md)\n\n<a id=\"validation\"></a>\n\n## ✅ Offline validation\n\nFrom a clone of this project, with Python 3.11+ and Git:\n\n```bash\npython scripts/run_harness_lab.py\n```\n\nThis offline example uses **scripted decisions and actual processes**. It exercises\nfailure → preparation → pause → controller restart → independent verification,\nwithout API calls, GPU use or model downloads. Inspect the printed `REPORT.json`\npath and its linked artifacts. Existing output is never overwritten; use\n`--output tmp/check-2` to repeat. It is not evidence of live-model capability.\n[Example source and checks](examples/harness-lab/README.md)\n\nRun the repository regression suite:\n\n```bash\npython scripts/run_all_tests.py\n```\n\nLatest local record (2026-09-07): **69/69 scripts passed in 156.0 s**.\nThe CI badge links to the current Windows, Linux and macOS results.\nLocal tests do not substitute for live-model or held-out evaluation.\n\nFor model comparisons, the [small paired-evaluation kit](docs/PAIRED_PILOT.md)\nprovides frozen tasks and actual grader-calibration logs. The six planned model\ntrials remain unrun; calibration is not evidence of skill uplift.\n\n[Controlled-trial checks](docs/CONTROLLED_TRIALS.md) add real failure/recovery\nlogs, restricted tools and unknown-usage stops, with scripted model responses.\nThe guide also provides a [bounded A/B command-line entrypoint](docs/CONTROLLED_TRIALS.md#one-bounded-ab-pair)\nfor model transport → reviewed commands → independent grading → sealed summary.\nLocal HTTP integration is tested; real-provider effectiveness is not yet measured.\n\n## Engineering and contributions\n\n[Engineering roadmap](docs/ENGINEERING_ROADMAP.md) · [Contributing](CONTRIBUTING.md) ·\n[Security and reporting](SECURITY.md) · [CI workflow](.github/workflows/validate.yml) ·\n[Reproduction feedback](https://github.com/lllllllama/RigorPilot-Skills/issues/new?template=reproduction.yml) ·\n[MIT license](LICENSE)\n\nKeep acceptance checks independent, retain failed evidence and review traces\nbefore publication. Do not publish credentials or unreviewed private repository data.\n[Agent guidance](AGENTS.md) · [Operating principles](references/agent-operating-principles.md) ·\n[Personalization policy](references/continuous-learning-policy.md)\n\n<details>\n<summary>Historical interface illustration—not execution evidence</summary>\n\n<img src=\"assets/annotated-readme-preview.png\" width=\"840\" alt=\"Historical MiniSeg interface illustration, not independently verified execution evidence\"/>\n\n[First attempt](examples/annotated-readme-demo/first-run/ANNOTATED_README.md) ·\n[After setup](examples/annotated-readme-demo/after-setup/ANNOTATED_README.md).\nThis older MiniSeg preview illustrates error, metric and authorization displays.\nIts execution provenance is not independently verified; it is excluded from benchmarks.\n\n</details>\n",
  "bytes": 14133,
  "sha": "47410a839ce27035b923d7eba5ccde7eb843fd7c179950056f25395de423ab7e",
  "repo_slug": "lllllllama/rigorpilot-skills",
  "fonte": "repo",
  "truncated": false,
  "api": "https://agentalog.com/api/listings/skl_lllllllama_rigorpilot_skills_safe_debug_bcba4fb8/readme"
}