{
  "markdown": "<p align=\"center\"><img src=\"./cover.png\" width=\"100%\" /></p>\n\n<h1 align=\"center\">autoconference-skill</h1>\n<p align=\"center\">\n  <em>Spawn a conference of autonomous researchers that compete, collaborate, and synthesize breakthroughs.</em>\n</p>\n<p align=\"center\">\n  <a href=\"#commands\">Commands</a> · <a href=\"#when-to-use\">When to Use</a> · <a href=\"#quick-start\">Quick Start</a> · <a href=\"#how-it-works\">How It Works</a> · <a href=\"#templates\">Templates</a> · <a href=\"#guide\">Guide</a> · <a href=\"./README-Ko-KR.md\">한국어</a>\n</p>\n<p align=\"center\">\n  <img src=\"https://img.shields.io/github/stars/wjgoarxiv/autoconference-skill?style=social\" />\n  <img src=\"https://img.shields.io/badge/license-MIT-blue\" />\n  <img src=\"https://img.shields.io/badge/python-3.8+-green\" />\n  <img src=\"https://img.shields.io/badge/version-2.0.0-orange\" />\n  <img src=\"https://img.shields.io/badge/skill-Claude%20Code%20%7C%20Codex%20%7C%20OpenCode%20%7C%20Gemini-blueviolet\" />\n</p>\n\n---\n\n> [!NOTE]\n> A Claude Code skill that orchestrates N parallel autoresearch agents in structured conference rounds -- with adversarial peer review and cross-researcher synthesis. Write a `conference.md` defining your research goal, and the conference handles hypothesis generation, experimentation, evaluation, and multi-agent iteration. Built on [autoresearch-skill](https://github.com/wjgoarxiv/autoresearch-skill). Works with Claude Code, Codex CLI, and Gemini CLI.\n\n### Example: sII Hydrate + Water .gro Generation\n\n| Round 1 (Naive) | Final (Converged) |\n|:---:|:---:|\n| ![Round 1](./examples/sii-hydrate-generation/snapshot_round1.png) | ![Final](./examples/sii-hydrate-generation/snapshot_final.png) |\n| Composite: 34.5 — water in slab | Composite: 99.9 — clean separation |\n\n| Convergence | Sub-metrics |\n|:---:|:---:|\n| ![Convergence](./examples/sii-hydrate-generation/plot_composite_convergence.png) | ![Sub-metrics](./examples/sii-hydrate-generation/plot_submetric_breakdown.png) |\n\n> 33 iterations across 3 researchers and 3 rounds. The conference learns to preserve crystal structure and exclude water from the hydrate slab. [See full example →](./examples/sii-hydrate-generation/)\n\n## Commands\n\nautoconference v2.0 provides 7 commands through a subcommand architecture:\n\n| Command | Description | Use Case |\n|---------|-------------|----------|\n| `/autoconference` | Core conference loop | Run N researchers through structured rounds with peer review |\n| `/autoconference:plan` | 8-step setup wizard | Interactively create a `conference.md` with dry-run evaluator gate |\n| `/autoconference:resume` | Checkpoint recovery | Resume an interrupted conference from last completed phase |\n| `/autoconference:analyze` | Post-conference analysis | Extract insights, failure modes, and transferable learnings |\n| `/autoconference:debate` | Adversarial debate | 2-researcher pro/con format with Opus judge |\n| `/autoconference:survey` | Literature survey | Multi-database systematic review with citation chains |\n| `/autoconference:ship` | Ship results | 8-phase pipeline to format results for publication |\n\n**Command chaining:**\n```\nplan ──> autoconference ──> ship              (standard pipeline)\nplan ──> autoconference ──> analyze ──> ship  (with post-analysis)\ndebate ──> autoconference                     (debate-informed experiment)\nsurvey ──> autoconference                     (literature-guided experiment)\nresume ──> autoconference                     (continuation)\n```\n\n## Features\n\n- **Multi-Agent Orchestration** -- N researchers explore different parts of the search space in parallel, then share findings after each round.\n- **Adversarial Peer Review** -- Opus-powered Reviewer agent challenges claims each round, catching overfitting and measurement noise before results propagate.\n- **Synthesis Over Selection** -- Final output combines complementary insights from multiple researchers, not just picks the winner.\n- **Dual Mode** -- Metric mode for numeric optimization, Qualitative mode for literature review and hypothesis generation.\n- **7 Subcommands** -- Plan, run, resume, analyze, debate, survey, and ship — chainable into full research pipelines.\n- **Guard Parameters** -- Conference-level safety constraints that override metric improvements when violated.\n- **Automatic Convergence** -- Detects plateau, budget exhaustion, or stall and triggers final synthesis automatically.\n- **Crash Recovery** -- 5-type recovery matrix for interrupted conferences (mid-research, mid-poster, mid-review, mid-transfer, pre-synthesis).\n- **Full Audit Trail** -- Per-researcher logs, poster sessions, peer reviews, conference-level TSV, and JSONL event stream.\n- **Built on autoresearch-skill** -- Each researcher runs the proven 5-stage experiment-evaluate-iterate loop.\n- **Safety Built In** -- Max iterations, time budgets, researcher timeouts, forbidden-change boundaries, and automatic rollback.\n\n## When to Use\n\nAutoconference builds on [autoresearch-skill](https://github.com/wjgoarxiv/autoresearch-skill)'s single-agent loop by adding parallel exploration, adversarial review, and cross-researcher synthesis. Use it when one agent exploring sequentially isn't enough -- either because the search space is too wide, or because self-evaluation alone can't be trusted.\n\n### autoconference vs alternatives\n\n|  | autoresearch-skill | Agent harness modes (team, /batch, ouroboros, etc.) | autoconference |\n|:---|:---|:---|:---|\n| **Agents** | 1 | N, each on a different subtask | N, all on the same problem with different strategies |\n| **Exploration** | Sequential strategy switching | Independent subtask decomposition | Search space partitioning + knowledge transfer between rounds |\n| **Validation** | Mechanical evaluator or agent self-judgment | Build/test pass | Opus adversarial reviewer (catches overfitting, noise, invalid claims) |\n| **Result integration** | Single best result | Per-subtask results merged | Synthesis -- combines complementary insights, not just picks a winner |\n| **Round structure** | None (continuous iteration) | None (one-shot dispatch) | Poster session → peer review → knowledge transfer per round |\n\n### Pick autoconference when\n\n- The **search space is wide enough to partition** -- N researchers exploring different regions in parallel cover more ground than one agent switching strategies sequentially\n- **Self-evaluation has blind spots** -- the adversarial Reviewer (Opus) catches overfitting, measurement noise, and Goodhart's Law effects that a single agent's evaluator misses\n- You need **synthesis, not selection** -- the final output should combine complementary findings from multiple approaches, not just take the best score\n- The research is **qualitative** (literature synthesis, hypothesis generation) and benefits from multiple perspectives converging on a shared taxonomy\n\n### Use autoresearch-skill instead when\n\n- The search space is **small enough for one agent** to cover within the iteration budget\n- You have a **mechanical evaluator** and trust its keep/revert decisions without external review\n- **Token cost matters** -- autoconference runs N researchers + reviewer + synthesizer per round, roughly N+2x the cost of a single autoresearch loop\n\n## Quick Start\n\n### 1. Copy-Paste Install\n\n> [!TIP]\n> Paste the block below directly into Claude Code. It clones the repo, installs the skill, and verifies the setup in one shot.\n\n```\nI want to install the autoconference-skill. Do these steps:\n1. git clone https://github.com/wjgoarxiv/autoconference-skill.git /tmp/autoconference-skill\n2. mkdir -p ~/.claude/skills/autoconference-skill && cp -r /tmp/autoconference-skill/SKILL.md /tmp/autoconference-skill/skills /tmp/autoconference-skill/scripts /tmp/autoconference-skill/assets /tmp/autoconference-skill/references /tmp/autoconference-skill/templates ~/.claude/skills/autoconference-skill/\n3. Test: python ~/.claude/skills/autoconference-skill/scripts/init_conference.py --goal \"test\" --metric \"score\" --direction minimize --researchers 2 --output /tmp/test-conference && echo \"OK: autoconference-skill installed\"\n4. Say \"autoconference-skill installed successfully\"\n```\n\n### 2. Manual Install\n\n```bash\ngit clone https://github.com/wjgoarxiv/autoconference-skill.git\ncd autoconference-skill\n\n# Symlink into your skills directory\nmkdir -p ~/.claude/skills\nln -s \"$(pwd)\" ~/.claude/skills/autoconference-skill\n```\n\n### 3. Other Tools\n\n| Tool | Install Command |\n|------|----------------|\n| Claude Code | Paste the copy-paste block above, or use the manual install |\n| Codex CLI | Symlink/copy the repository directory so `SKILL.md`, `skills/`, `references/`, and `scripts/` stay together |\n| Gemini CLI | Symlink/copy the repository directory and keep `gemini-extension.json` at the repo root |\n\n### Other Platforms\n\n| Platform | Skills Path | Install Command |\n|----------|-------------|-----------------|\n| **Claude Code** | `~/.claude/skills/autoconference-skill/` | See above |\n| **Codex CLI** | `~/.codex/skills/autoconference-skill/` | `mkdir -p ~/.codex/skills && ln -s \"$(pwd)\" ~/.codex/skills/autoconference-skill` |\n| **OpenCode** | `~/.config/opencode/skills/autoconference-skill/` | `mkdir -p ~/.config/opencode/skills && ln -s \"$(pwd)\" ~/.config/opencode/skills/autoconference-skill` |\n| **Gemini CLI** | `~/.gemini/skills/autoconference-skill/` | `mkdir -p ~/.gemini/skills && ln -s \"$(pwd)\" ~/.gemini/skills/autoconference-skill` |\n\n## Usage\n\n### Prompt Optimization Tournament\n\nThree researchers compete on prompt accuracy — each specializing in instruction phrasing, few-shot selection, and chain-of-thought formatting.\n\n```\nRun an autoconference using templates/prompt-optimization.md.\nGoal: maximize accuracy on my classification benchmark.\n3 researchers, 3 rounds.\n```\n\n### Code Performance Competition\n\nAlgorithmic, data-structure, and low-level researchers independently optimize the same codebase, then cross-pollinate validated wins each round.\n\n```\nRun an autoconference using templates/code-performance.md.\nMetric: wall-clock time on my benchmark suite.\nDirection: minimize. Target: < 200ms.\n```\n\n### Literature Synthesis Conference\n\nQualitative mode. Three researchers survey the same topic from foundational, recent, and cross-domain angles, then synthesize a unified taxonomy.\n\n```\nRun an autoconference in qualitative mode.\nGoal: survey LLM agent papers from 2022-2025.\n3 researchers, 2 rounds. Synthesize findings into a taxonomy.\n```\n\n### Scaffold a New Conference\n\n```bash\npython scripts/init_conference.py \\\n  --goal \"Optimize inference latency\" \\\n  --metric \"p95_latency_ms\" \\\n  --direction minimize \\\n  --target \"< 50\" \\\n  --researchers 3 \\\n  --devils-advocate yes \\\n  --strategy assigned \\\n  --output ./latency-conference/\n```\n\nThen edit the generated `conference.md` to fill in your `Current Approach`, `Search Space`, and researcher focus areas. When ready:\n\n```\nRun the autoconference on my conference.md\n```\n\nClaude loads `SKILL.md`, reads `conference.md`, and orchestrates the full conference -- all rounds, peer review, and final synthesis.\nBefore Phase 1, it must summarize researcher count, iterations/budget, success definition, and Critic/Devil's Advocate setting, then wait for your final confirmation.\n\n## How It Works\n\n```\n+----------------------------------------------------------+\n|                     CONFERENCE ROUND                     |\n|                                                          |\n|  Phase 1: INDEPENDENT RESEARCH (parallel)               |\n|  +----------+  +----------+  +----------+               |\n|  |Researcher|  |Researcher|  |Researcher|  Each runs N  |\n|  |    A     |  |    B     |  |    C     |  autoresearch |\n|  | (iter x N)|  | (iter x N)|  | (iter x N)|  iterations |\n|  +----+-----+  +----+-----+  +----+-----+               |\n|       |              |             |                     |\n|  Phase 2: POSTER SESSION                                 |\n|  +----------------------------------------------+       |\n|  | Session Chair collects all logs,             |       |\n|  | surfaces key findings & deltas               |       |\n|  +----------------------+-----------------------+       |\n|                         |                               |\n|  Phase 3: PEER REVIEW (adversarial)                     |\n|  +----------------------------------------------+       |\n|  | Reviewer agent challenges claims:            |       |\n|  | - \"Did metric actually improve?\"             |       |\n|  | - \"Is this overfitting?\"                     |       |\n|  | - \"Could this be measurement noise?\"         |       |\n|  +----------------------+-----------------------+       |\n|                         |                               |\n|  Phase 4: KNOWLEDGE TRANSFER                            |\n|  +----------------------------------------------+       |\n|  | Validated findings shared back to            |       |\n|  | all researchers for next round               |       |\n|  +----------------------------------------------+       |\n|                                                          |\n+----------------------------------------------------------+\n          |\n          v  Convergence check -> next round or final synthesis\n```\n\n## The `conference.md` Format\n\n| Section | Purpose |\n|---------|---------|\n| `Goal` | What the conference should achieve |\n| `Mode` | `metric` (numeric optimization) or `qualitative` (reasoning quality) |\n| `Success Metric` | Metric name, target, direction (metric mode only) |\n| `Success Criteria` | Natural language description of \"good\" (qualitative mode only) |\n| `Researchers` | Count, iterations per round, max rounds |\n| `Pre-Flight Gate` | Researcher count, calculated iteration budget, success definition, Critic/Devil's Advocate setting, and pending final confirmation |\n| `Search Space` | What researchers can and cannot modify |\n| `Search Space Partitioning` | `assigned` (each researcher has a focus) or `free` (overlap allowed) |\n| `Constraints` | Max iterations, time budget, researcher timeout |\n| `Current Approach` | Baseline description |\n| `Shared Knowledge` | Auto-populated after each round with validated findings |\n| `Conference Log` | Auto-maintained round-by-round history |\n\nSee `assets/conference_template.md` for the full template.\n\n## Agent Roles\n\n| Role | Model | Count | Responsibility |\n|------|-------|-------|----------------|\n| **Conference Chair** | Sonnet | 1 | Orchestrator -- manages rounds, spawns researchers, detects convergence, triggers synthesis |\n| **Researcher** | Sonnet | N | Runs the autoresearch 5-stage loop within assigned search space |\n| **Session Chair** | Haiku | 1 | Lightweight summarizer -- collects logs and produces poster session summary after each round |\n| **Reviewer** | Opus | 1 | Adversarial critic -- challenges claims, checks for overfitting/noise, assigns verdicts |\n| **Synthesizer** | Opus | 1 | Runs once at end -- combines complementary insights from all researchers |\n\n## Templates\n\nReady-to-use `conference.md` configs for common tasks:\n\n| Template | Mode | Use Case |\n|----------|------|----------|\n| `templates/quick-conference.md` | metric | 2 researchers, 2 rounds -- test if your problem benefits from the conference format |\n| `templates/prompt-optimization.md` | metric | Optimize LLM prompt accuracy with 3 specialized researchers |\n| `templates/code-performance.md` | metric | Optimize code speed with algorithmic, data-structure, and low-level researchers |\n| `templates/research-synthesis.md` | qualitative | Literature exploration across foundational, recent, and cross-domain angles |\n| `templates/debate-mode.md` | qualitative | 2-researcher adversarial debate with structured rounds and Opus judge |\n| `templates/survey-mode.md` | qualitative | Multi-database literature survey with citation chain tracking |\n\n## Configuration Options\n\n| Field | Default | Description |\n|-------|---------|-------------|\n| `mode` | `metric` | `metric` or `qualitative` |\n| `count` | -- | Number of researcher agents |\n| `iterations_per_round` | 5 | Autoresearch iterations each researcher runs per round |\n| `max_rounds` | 4 | Maximum conference rounds before forced synthesis |\n| `max_total_iterations` | -- | Hard cap across all researchers and rounds |\n| `time_budget` | -- | Wall-clock limit for the entire conference |\n| `researcher_timeout` | -- | Per-researcher timeout per round |\n| `strategy` | `free` | `assigned` (focus areas) or `free` (open exploration) |\n| `guard` | -- | Safety constraint enforced on ALL researchers (violations revert regardless of metric) |\n| `noise_runs` | 1 | Repeated evaluations to average for noise reduction |\n| `min_consensus_delta` | 0 | Minimum average improvement across kept researchers to advance |\n\n## What You Get\n\nThe conference writes both human-readable reports and machine-readable logs. The sII hydrate example shows the same contract in a completed run.\n\n| Output | Purpose | Example artifact | How to inspect |\n|--------|---------|------------------|----------------|\n| `conference.md` | Input config plus round log and shared knowledge | [`examples/sii-hydrate-generation/conference.md`](./examples/sii-hydrate-generation/conference.md) | Read first to confirm goal, metric, guardrails |\n| `conference_results.tsv` | Master table: all rounds, researchers, metrics, verdicts | [`conference_results.tsv`](./examples/sii-hydrate-generation/conference_results.tsv) | Run `python scripts/validate_package.py` or import into a spreadsheet |\n| `researcher_A_results.tsv` | Per-researcher TSV compatible with single-agent result logs | [`researcher_A_results.tsv`](./examples/sii-hydrate-generation/researcher_A_results.tsv) | Check 8-column schema from `references/results-logging.md` |\n| `conference_events.jsonl` | Append-only event stream for resume/progress tools | [`conference_events.jsonl`](./examples/sii-hydrate-generation/conference_events.jsonl) | One JSON object per line |\n| `poster_session_round_N.md` | Session-chair summary after each round | [`poster_session_round_1.md`](./examples/sii-hydrate-generation/poster_session_round_1.md) | Compare researcher findings |\n| `peer_review_round_N.md` | Adversarial validation of claims | [`peer_review_round_1.md`](./examples/sii-hydrate-generation/peer_review_round_1.md) | Look for challenged/overturned claims |\n| `synthesis.md` | Final cross-researcher synthesis | [`synthesis.md`](./examples/sii-hydrate-generation/synthesis.md) | Read before the full report |\n| `final_report.md` | Executive summary and evidence trail | [`final_report.md`](./examples/sii-hydrate-generation/final_report.md) | Shareable result summary |\n\nQuick validation from the repo root:\n\n```bash\npython scripts/validate_package.py\nbash scripts/check_conference.sh examples/sii-hydrate-generation\n```\n\n## Output Files\n\n| File | Description |\n|------|-------------|\n| `conference.md` | User config (updated with log entries each round) |\n| `conference_results.tsv` | Master conference-level TSV with all iterations and peer review verdicts |\n| `researcher_A_log.md` | Detailed per-researcher iteration log |\n| `researcher_A_results.tsv` | Per-researcher TSV (same format as autoresearch) |\n| `poster_session_round_N.md` | Session Chair summary for each round |\n| `peer_review_round_N.md` | Reviewer verdicts for each round |\n| `synthesis.md` | Final synthesized output from Synthesizer |\n| `final_report.md` | Executive summary with full conference history |\n\n## Overnight Runs\n\nTo run a conference overnight, use the universal loop script:\n\n```bash\n# Option A: Foreground (simplest)\nbash scripts/autoconference-loop.sh ./my-conference/\n\n# Option B: Background with nohup (no tmux needed)\nnohup bash scripts/autoconference-loop.sh ./my-conference/ > conference.log 2>&1 &\n\n# Option C: Background with tmux (best experience)\ntmux new-session -d -s conference 'bash scripts/autoconference-loop.sh ./my-conference/'\n\n# Check progress anytime\nbash scripts/check_conference.sh ./my-conference/\n```\n\nThe script auto-detects your CLI tool, handles round restarts, and checks for conference completion. Works with Claude Code, Codex CLI, OpenCode, and Gemini CLI.\n\n## Relationship to autoresearch-skill\n\nEach researcher in a conference runs the **autoresearch loop** -- the same autonomous experiment-evaluate-iterate cycle from [autoresearch-skill](https://github.com/wjgoarxiv/autoresearch-skill). Autoconference adds three layers on top:\n\n1. **Multi-agent orchestration** -- N researchers explore different parts of the search space in parallel\n2. **Adversarial peer review** -- A Reviewer agent challenges findings each round (catches what self-evaluation misses)\n3. **Synthesis** -- A Synthesizer combines complementary insights rather than just picking the best result\n\nUse autoresearch-skill for a single focused research loop. Use autoconference when your search space is large enough to partition, when diversity of approach matters, or when you want external validation of results.\n\n## Guide\n\nComprehensive documentation is available in the `guide/` directory:\n\n| Guide | Topic |\n|-------|-------|\n| [Getting Started](guide/getting-started.md) | 60-second quickstart + domain cheat sheet |\n| [Core Conference](guide/autoconference.md) | 4-phase round structure, convergence, overnight runs |\n| [Plan](guide/plan.md) | 8-step setup wizard |\n| [Resume](guide/resume.md) | Checkpoint recovery |\n| [Analyze](guide/analyze.md) | Post-conference insight extraction |\n| [Debate](guide/debate.md) | Adversarial 2-researcher format |\n| [Survey](guide/survey.md) | Multi-database literature survey |\n| [Ship](guide/ship.md) | Conference results to publication |\n| [Chains](guide/chains-and-combinations.md) | Command chaining patterns |\n| [Advanced](guide/advanced-patterns.md) | Guards, noise, worktrees, CI/CD |\n| [Troubleshooting](guide/troubleshooting.md) | Common failure modes and fixes |\n| [Self-Improvement](guide/self-improvement.md) | Convert confusing runs into safe improvement plans and eval cases |\n| [Eval Coverage](evals/README.md) | Routing and safety scenario categories |\n\nSee also: [COMPARISON.md](./COMPARISON.md) for autoconference vs alternatives.\n\n## Cross-Platform Compatibility\n\n| Platform | Status | Install |\n|----------|--------|---------|\n| Claude Code | Ready | Plugin install or manual symlink |\n| Codex CLI | Ready | See `.codex/INSTALL.md` |\n| OpenCode | Ready | See `.opencode/INSTALL.md` |\n| Gemini CLI | Ready | Uses `gemini-extension.json` auto-discovery |\n\n## Requirements\n\n| Requirement | Details |\n|-------------|---------|\n| **Python** | 3.8+ (stdlib only — for `scripts/init_conference.py` only) |\n| **LLM CLI** | Any CLI with subagent support (Claude Code, Codex CLI, OpenCode, Gemini CLI) |\n| **autoresearch-skill** | Referenced by each researcher agent's prompt |\n\n## Contributing\n\nSee [CONTRIBUTING.md](./CONTRIBUTING.md) for detailed guidelines.\n\n1. Fork the repository\n2. Create a feature branch (`git checkout -b feature/your-feature`)\n3. Commit your changes (`git commit -m 'Add your feature'`)\n4. Push to the branch (`git push origin feature/your-feature`)\n5. Open a pull request\n\n## License\n\nMIT -- see [LICENSE](./LICENSE) for details.\n",
  "bytes": 23110,
  "sha": "0fde3e68723e8b11c6b7e8a6184f9a4c3d642ede583d4814496cfc3228e9850c",
  "repo_slug": "wjgoarxiv/autoconference-skill",
  "fonte": "repo",
  "truncated": false,
  "api": "https://agentalog.com/api/listings/plg_wjgoarxiv_autoconference_skill_f6e90b25/readme"
}