{
  "markdown": "# Research Harness\n\n[![Repository verification](https://github.com/WILLOSCAR/research-units-pipeline-skills/actions/workflows/verify.yml/badge.svg)](https://github.com/WILLOSCAR/research-units-pipeline-skills/actions/workflows/verify.yml)\n\n**Research should leave a trail, not just an answer.**\n\nA long research task can produce a polished PDF and still leave basic questions\nunanswered: Which sources support this paragraph? What changed after the last\nfailure? Can the work resume tomorrow without reconstructing a chat? What did\n`PASS` actually verify?\n\nResearch Harness turns a research goal into a file-first, recoverable Run. It\norganizes focused Skills into explicit Workflows, preserves intermediate\nArtifacts and decisions, checks observable contracts, and points failures back\nto the smallest repair surface.\n\n```text\nGoal -> Run -> Evidence -> Artifact\n```\n\nIt is not an autonomous-scientist claim. It is infrastructure for making\nagent-assisted research inspectable, resumable, and honest about what has—and\nhas not—been proven.\n\n## See A Run In Five Minutes\n\nResearch Harness currently runs from a source checkout with Python 3.10+ and\n[uv](https://docs.astral.sh/uv/):\n\n```bash\ngit clone https://github.com/WILLOSCAR/research-units-pipeline-skills.git\ncd research-units-pipeline-skills\nuv sync --locked\n\nuv run rh goal create \\\n  --goal \"Understand test-time adaptation for robotics and decide what to read\" \\\n  --workflow research-brief \\\n  --workspace workspaces/robot-adaptation\n\nuv run rh run start --workspace workspaces/robot-adaptation\n```\n\nThe Run advances until it finishes or reaches an unmet prerequisite. For\n`research-brief`, inspect the paper set, taxonomy, outline, and C2 review block,\nthen continue:\n\n```bash\nuv run rh run status --workspace workspaces/robot-adaptation\nuv run rh run approve --workspace workspaces/robot-adaptation --checkpoint C2\nuv run rh run resume --workspace workspaces/robot-adaptation\nuv run rh evidence inspect --workspace workspaces/robot-adaptation --excerpt\n```\n\nThe Workspace now contains the readable deliverable and its evidence trail:\n\n```text\nGOAL.md                  requested outcome and constraints\nUNITS.csv                explicit plan and current Unit state\nDECISIONS.md             human checkpoints and choices\npapers/ + outline/       research evidence and intermediate structure\noutput/                  deliverable, scorecards, audits, repair reports\n.harness/                Run identity, Attempts, Events, hashes, provenance\n```\n\nIf a contract fails, ask the Harness where repair belongs:\n\n```bash\nuv run rh improve diagnose --workspace workspaces/robot-adaptation\n```\n\n## Choose The Deliverable\n\nUsers choose a Workflow by outcome; Skills and Units stay implementation\ndetails until inspection or repair is necessary.\n\n| You want to… | Workflow | Required starting point | Main deliverable |\n|---|---|---|---|\n| Understand a topic and decide what to read | `research-brief` | topic | `output/SNAPSHOT.md` |\n| Review one paper or manuscript | `paper-review` | manuscript | `output/REVIEW.md` |\n| Synthesize studies under an approved protocol | `evidence-review` | review question | `output/SYNTHESIS.md` |\n| Write a literature survey or bounded report | `arxiv-survey` | topic and delivery constraints | `output/DRAFT.md` |\n| Deliver that Survey as LaTeX and PDF | `arxiv-survey-latex` | topic and delivery constraints | `latex/main.pdf` |\n| Develop literature-grounded research directions | `idea-brainstorm` | topic and scope | `output/REPORT.md` |\n| Turn a fixed source set into a tutorial | `source-tutorial` | source pack and audience | tutorial, article PDF, slides |\n\nIn Codex or Claude Code, the activation surface is deliberately one sentence:\n\n```text\nUse research-brief to map test-time adaptation for robotics and tell me what to read first.\nUse paper-review to review the attached manuscript and trace every major concern to the paper.\nUse arxiv-survey-latex to write an 8-10 page course paper on RAG evaluation and produce a PDF.\nUse source-tutorial to turn sources/manifest.yml into a tutorial for senior software engineers.\n```\n\n`graduate-paper` remains a research-stage Chinese thesis path, not one of the\nseven executable Pipeline contracts.\n\nInput boundaries are intentional. `paper-review` will not invent a manuscript;\n`source-tutorial` will not invent a source pack; `evidence-review` writes a\nprotocol and pauses for approval before retrieval. See the\n[usage guides](readme/README.en.md) for those setup paths.\n\n## What Changes When Research Becomes A Run\n\nWithout a Harness, a research agent usually leaves a final answer and a long\nconversation. With Research Harness, each transition has an inspectable owner:\n\n```mermaid\nflowchart LR\n    G[\"Goal\"] --> W[\"Workflow\"]\n    W --> P[\"Pinned Pipeline contract\"]\n    P --> U[\"Recoverable Units\"]\n    U --> A[\"Research Artifacts\"]\n    A --> C[\"Completion checks\"]\n    C --> E[\"Run Evidence\"]\n    E --> D[\"Bounded diagnosis\"]\n    D -. \"repair and rerun\" .-> U\n```\n\nThree mechanisms make that trail useful:\n\n1. **The contract is pinned.** `harness-lock.v2` snapshots the selected Pipeline\n   and hashes its inheritance bundle, Skill implementations, and Harness Kernel.\n   An active Run fails closed if the Pipeline or Kernel drifts; it cannot silently\n   continue under different rules.\n2. **Completion is evidence-backed.** A `DONE` cell alone is not success. The\n   Attempt, required outputs, Artifact hashes, Workflow checks, Manifest, and\n   Completion Event must agree.\n3. **Failure has an address.** Doctor, Audit, scorecards, and the Failure ledger\n   distinguish an observable defect from its owning repair surface. Improvement\n   diagnoses; it does not rewrite the Harness in place.\n\nHuman checkpoints use the same discipline. Approval is bound to the reviewed\nArtifact hashes, so changing an approved outline, scope, or protocol revokes the\nstale authorization.\n\n## What A PASS Means\n\nResearch Harness separates three claims that are easy to blur:\n\n| Layer | A PASS establishes | It does not establish |\n|---|---|---|\n| Execution integrity | Attempts, state, Manifests, hashes, and provenance agree | that the answer is good |\n| Contract acceptance | required Artifacts satisfy observable Workflow checks | scientific truth or exhaustive retrieval |\n| Research quality | usefulness and correctness on realistic inputs | validity beyond the evaluated cases |\n\nThe repository implements the first two layers. The third needs repeated Runs,\nheld-out evaluation, and expert judgment. Reports use qualified evidence rather\nthan turning every green check into a research-quality claim.\n\n## The Survey Failure That Shaped The Gate\n\nThe Survey writer can bootstrap provisional prose from structured evidence packs\nand versioned templates. Early versions completed the delivery path but left too\nmuch of that scaffold in the paper: the historical course-paper sample matches\ntemplate fragments in **96/140 sentences (68.6%)**.\n\nThat failure is now a contract, not a warning:\n\n- `front-matter-writer` checks the abstract, introduction, related work,\n  discussion, and conclusion before merge;\n- `subsection-writer` and `writer-selfloop` check H3 prose;\n- `pipeline-auditor` checks the whole merged draft, selected asset hashes, and\n  the three template-owning Skill implementations;\n- pipeline voice such as “this run” is blocking reader-facing residue;\n- the whole-draft limit is <=10%.\n\nThe current published replay completes all 49 Units under the current contract:\n\n| Evidence | Result |\n|---|---:|\n| Required Workflow checks | 31/31 PASS |\n| Target Artifacts | 75/75 present |\n| Harness Kernel lock | 35/35 matched |\n| Ledger integrity issues | 0 |\n| Template residue | 0/226 sentences (0.0%) |\n| PDF delivery | 10 pages |\n\nThis proves attainability for one retained Artifact set. It does not prove\nauthorship, semantic originality, autonomous generation, cross-topic\ncalibration, or expert paper quality. The Run used manual Artifact revalidation\nand a dirty worktree; a clean, from-scratch reproduction remains open. Inspect\nthe [current-contract evidence](examples/course-paper-residue-pass/README.md)\nand the [historical failure baseline](examples/course-paper-pilot/README.md).\n\n## Published Evidence\n\nThe repository publishes curated evidence rather than private Workspaces:\n\n| Snapshot | What it demonstrates | Boundary |\n|---|---|---|\n| [`course-paper-residue-pass`](examples/course-paper-residue-pass/README.md) | current v2 contract acceptance, 0/226 residue, 10-page PDF | manual replay, dirty revision, one topic |\n| [`course-paper-pilot`](examples/course-paper-pilot/README.md) | completed delivery and a reproducible 68.6% failure baseline | historical contract; fails the current writing gate |\n| [`research-brief-real-source-proof`](examples/research-brief-real-source-proof/README.md) | one live-arXiv briefing delivery | historical v1 protocol, one topic |\n| [`research-brief-harness-proof`](examples/research-brief-harness-proof/README.md) | deterministic recovery and Audit evidence | synthetic sources, historical v1 protocol |\n\nScorecard fixtures and failure-repair regressions cover `paper-review`,\n`idea-brainstorm`, `evidence-review`, and `source-tutorial`. Cross-topic\nstability, measured model-token benchmarks, expert comparison, and automatic\nHarness-candidate promotion remain open.\n\n## Runtime Requirements\n\n- Python 3.10+ and `uv` for the CLI;\n- `pdftotext` for Source Tutorial PDF ingestion;\n- `latexmk`, XeLaTeX, BibTeX, and `pdfinfo` for LaTeX/PDF delivery.\n\nThe Python package declares `PyYAML` and `pypdf`; maintainer dependencies are in\nthe `test` extra. GitHub Actions installs the same TeX/Poppler boundary used by\nthe PDF tests.\n\n## Maintainer Verification\n\nRun the same checks as `.github/workflows/verify.yml`:\n\n```bash\nuv run --locked python scripts/validate_repo.py --strict\nuv run --locked python scripts/readiness_audit.py --strict\nuv run --locked python scripts/audit_skills.py --fail-on WARN\nuv run --locked python scripts/audit_workflow_context.py\nuv run --locked --extra test ruff check .\nuv run --locked --extra test python -m pytest -q\n```\n\nWhen extending a Workflow, keep its Pipeline contract, Unit template, owned\nSkills, tests, and evidence claim aligned. Do not raise a proof state without a\ncompleted Run or a failure-repair regression that supports it.\n\n## Documentation\n\n- [Architecture](docs/AUTO_RESEARCH_DESIGN_SYSTEM.md)\n- [Workflow catalog and proof states](docs/PIPELINE_TAXONOMY.md)\n- [Canonical product glossary](CONTEXT.md)\n- [Implementation-language map](docs/PROJECT_LANGUAGE.md)\n- [Roadmap](docs/HARNESS_ROADMAP.md)\n- [Current readiness](docs/HARNESS_READINESS.md)\n- [Schemas](docs/SCHEMAS.md)\n- [Architecture decisions](docs/adr/)\n- [Detailed usage guides](readme/README.en.md)\n\n[中文 README](README.zh-CN.md)\n\n## Star History\n\n<a href=\"https://www.star-history.com/?repos=WILLOSCAR%2Fresearch-units-pipeline-skills&type=date&legend=top-left\">\n  <picture>\n    <source media=\"(prefers-color-scheme: dark)\" srcset=\"assets/star-history/star-history-dark.svg\">\n    <img alt=\"Star history chart\" src=\"assets/star-history/star-history-light.svg\">\n  </picture>\n</a>\n",
  "bytes": 11181,
  "sha": "717ffa9f7375f36d3adb3c946f7c27f797a6887197922ee84728b57ad473af55",
  "repo_slug": "willoscar/research-units-pipeline-skills",
  "fonte": "repo",
  "truncated": false,
  "api": "https://api.agentalog.com/api/listings/skl_willoscar_research_units_pipeline_skills_a379f914/readme"
}