{
  "markdown": "# MCP as a Judge ⚖️\n\nmcp-name: io.github.OtherVibes/mcp-as-a-judge\n\n<div align=\"left\">\n  <img src=\"assets/mcp-as-a-judge.png\" alt=\"MCP as a Judge Logo\" width=\"200\">\n</div>\n\n> MCP as a Judge acts as a validation layer between AI coding assistants and LLMs, helping ensure safer and higher-quality code.\n\n\n[![License: MIT](https://img.shields.io/badge/License-MIT-green.svg)](https://opensource.org/license/mit/)\n[![Python 3.13+](https://img.shields.io/badge/python-3.13+-blue.svg)](https://www.python.org/downloads/)\n[![MCP Compatible](https://img.shields.io/badge/MCP-Compatible-green.svg)](https://modelcontextprotocol.io/)\n\n[![CI](https://github.com/OtherVibes/mcp-as-a-judge/workflows/CI/badge.svg)](https://github.com/OtherVibes/mcp-as-a-judge/actions/workflows/ci.yml)\n[![Release](https://github.com/OtherVibes/mcp-as-a-judge/workflows/Release/badge.svg)](https://github.com/OtherVibes/mcp-as-a-judge/actions/workflows/release.yml)\n[![PyPI version](https://img.shields.io/pypi/v/mcp-as-a-judge.svg)](https://pypi.org/project/mcp-as-a-judge/)\n\n\n\n**MCP as a Judge** is a **behavioral MCP** that strengthens AI coding assistants by requiring explicit LLM evaluations for:\n- Research, system design, and planning\n- Code changes, testing, and task-completion verification\n\nIt enforces evidence-based research, reuse over reinvention, and human-in-the-loop decisions.\n\n> If your IDE has rules/agents (Copilot, Cursor, Claude Code), keep using them—this Judge adds enforceable approval gates on plan, code diffs, and tests.\n\n\n## Key problems with AI coding assistants and LLMs\n- Treat LLM output as ground truth; skip research and use outdated information\n- Reinvent the wheel instead of reusing libraries and existing code\n- Cut corners: code below engineering standards and weak tests\n- Make unilateral decisions when requirements are ambiguous or plans change\n- Security blind spots: missing input validation, injection risks/attack vectors, least‑privilege violations, and weak defensive programming\n\n\n## **Vibe coding doesn’t have to be frustrating**\n\n### What it enforces\n- Evidence‑based research and reuse (best practices, libraries, existing code)\n- Plan‑first delivery aligned to user requirements\n- Human‑in‑the‑loop decisions for ambiguity and blockers\n- Quality gates on code and tests (security, performance, maintainability)\n\n### Key capabilities\n- Intelligent code evaluation via MCP [sampling](https://modelcontextprotocol.io/docs/learn/client-concepts#sampling); enforces software‑engineering standards and flags security/performance/maintainability risks\n- Comprehensive plan/design review: validates architecture, research depth, requirements fit, and implementation approach\n- User‑driven decisions via MCP [elicitation](https://modelcontextprotocol.io/docs/learn/client-concepts#elicitation): clarifies requirements, resolves obstacles, and keeps choices transparent\n- Security validation in system design and code changes\n\n\n\n### Tools and how they help\n| Tool | What it solves |\n|------|-----------------|\n| `set_coding_task` | Creates/updates task metadata; classifies task_size; returns next-step workflow guidance |\n| `get_current_coding_task` | Recovers the latest task_id and metadata to resume work safely |\n| `judge_coding_plan` | Validates plan/design; requires library selection and internal reuse maps; flags risks |\n| `judge_code_change` | Reviews unified Git diffs for correctness, reuse, security, and code quality |\n| `judge_testing_implementation` | Validates tests using real runner output and optional coverage |\n| `judge_coding_task_completion` | Final gate ensuring plan, code, and tests approvals before completion |\n| `raise_missing_requirements` | Elicits missing details and decisions to unblock progress |\n| `raise_obstacle` | Engages the user on trade‑offs, constraints, and enforced changes |\n\n## 🚀 **Quick Start**\n\n### **Requirements & Recommendations**\n\n#### **MCP Client Prerequisites**\n\nMCP as a Judge is heavily dependent on **MCP Sampling** and **MCP Elicitation** features for its core functionality:\n\n- **[MCP Sampling](https://modelcontextprotocol.io/docs/learn/client-concepts#sampling)** - Required for AI-powered code evaluation and judgment\n- **[MCP Elicitation](https://modelcontextprotocol.io/docs/learn/client-concepts#elicitation)** - Required for interactive user decision prompts\n\n#### **System Prerequisites**\n\n- **Docker Desktop** / **Python 3.13+** - Required for running the MCP server\n\n#### **Supported AI Assistants**\n\n| AI Assistant | Platform | MCP Support | Status | Notes |\n|---------------|----------|-------------|---------|-------|\n| **GitHub Copilot** | Visual Studio Code | ✅ Full | **Recommended** | Complete MCP integration with sampling and elicitation |\n| **Claude Code** | - | ⚠️ Partial | Requires LLM API key | [Sampling Support feature request](https://github.com/anthropics/claude-code/issues/1785)<br>[Elicitation Support feature request](https://github.com/anthropics/claude-code/issues/2799) |\n| **Cursor** | - | ⚠️ Partial | Requires LLM API key | MCP support available, but sampling/elicitation limited |\n| **Augment** | - | ⚠️ Partial | Requires LLM API key | MCP support available, but sampling/elicitation limited |\n| **Qodo** | - | ⚠️ Partial | Requires LLM API key | MCP support available, but sampling/elicitation limited |\n\n**✅ Recommended setup:** GitHub Copilot + VS Code — full MCP sampling; no API key needed.\n\n**⚠️ Critical:** For assistants without full MCP sampling (Cursor, Claude Code, Augment, Qodo), you MUST set `LLM_API_KEY`. Without it, the server cannot evaluate plans or code. See [LLM API Configuration](#-llm-api-configuration-optional).\n\n**💡 Tip:** Prefer large context models (≥ 1M tokens) for better analysis and judgments.\n\n### If the MCP server isn’t auto‑used\nFor troubleshooting, visit the [FAQs section](#faq).\n\n## 🔧 **MCP Configuration**\n\nConfigure **MCP as a Judge** in your MCP-enabled client:\n\n### **Method 1: Using Docker (Recommended)**\n\n#### One‑click install for VS Code (MCP)\n\n[![Install for MCP as a Judge](https://img.shields.io/badge/VS_Code-Install_for_MCP_as_a_Judge-0098FF?style=flat-square&logo=visualstudiocode&logoColor=white)](https://insiders.vscode.dev/redirect/mcp/install?name=mcp-as-a-judge&inputs=%5B%5D&config=%7B%22command%22%3A%22docker%22%2C%22args%22%3A%5B%22run%22%2C%22-i%22%2C%22--rm%22%2C%22--pull%3Dalways%22%2C%22ghcr.io%2Fothervibes%2Fmcp-as-a-judge%3Alatest%22%5D%7D)\n\n\n\nNotes:\n- VS Code controls the sampling model; select it via “MCP: List Servers → mcp-as-a-judge → Configure Model Access”.\n\n\n1. **Configure MCP Settings:**\n\n   Add this to your MCP client configuration file:\n\n   ```json\n   {\n     \"command\": \"docker\",\n     \"args\": [\"run\", \"--rm\", \"-i\", \"--pull=always\", \"ghcr.io/othervibes/mcp-as-a-judge:latest\"],\n     \"env\": {\n       \"LLM_API_KEY\": \"your-openai-api-key-here\",\n       \"LLM_MODEL_NAME\": \"gpt-4o-mini\"\n     }\n   }\n   ```\n\n   **📝 Configuration Options (All Optional):**\n   - **LLM_API_KEY**: Optional for GitHub Copilot + VS Code (has built-in MCP sampling)\n   - **LLM_MODEL_NAME**: Optional custom model (see [Supported LLM Providers](#supported-llm-providers) for defaults)\n   - The `--pull=always` flag ensures you always get the latest version automatically\n\n   Then manually update when needed:\n\n   ```bash\n   # Pull the latest version\n   docker pull ghcr.io/othervibes/mcp-as-a-judge:latest\n   ```\n\n### **Method 2: Using uv**\n\n1. **Install the package:**\n\n   ```bash\n   uv tool install mcp-as-a-judge\n   ```\n\n2. **Configure MCP Settings:**\n\n   The MCP server may be automatically detected by your MCP‑enabled client.\n\n   **📝 Notes:**\n   - **No additional configuration needed for GitHub Copilot + VS Code** (has built-in MCP sampling)\n   - LLM_API_KEY is optional and can be set via environment variable if needed\n\n3. **To update to the latest version:**\n\n   ```bash\n   # Update MCP as a Judge to the latest version\n   uv tool upgrade mcp-as-a-judge\n   ```\n### Select a sampling model in VS Code\n- Open Command Palette (Cmd/Ctrl+Shift+P) → “MCP: List Servers”\n- Select the configured server “mcp-as-a-judge”\n- Choose “Configure Model Access”\n- Check your preferred model(s) to enable sampling\n\n\n\n## 🔑 **LLM API Configuration (Optional)**\n\nFor [AI assistants without full MCP sampling support](#supported-ai-assistants) you can configure an LLM API key as a fallback. This ensures MCP as a Judge works even when the client doesn't support MCP sampling.\n\n- Set `LLM_API_KEY` (unified key). Vendor is auto-detected; optionally set `LLM_MODEL_NAME` to override the default.\n\n### **Supported LLM Providers**\n\n| Rank | Provider | API Key Format | Default Model | Notes |\n|------|----------|----------------|---------------|-------|\n| **1** | **OpenAI** | `sk-...` | `gpt-4.1` | Fast and reliable model optimized for speed |\n| **2** | **Anthropic** | `sk-ant-...` | `claude-sonnet-4-20250514` | High-performance with exceptional reasoning |\n| **3** | **Google** | `AIza...` | `gemini-2.5-pro` | Most advanced model with built-in thinking |\n| **4** | **Azure OpenAI** | `[a-f0-9]{32}` | `gpt-4.1` | Same as OpenAI but via Azure |\n| **5** | **AWS Bedrock** | AWS credentials | `anthropic.claude-sonnet-4-20250514-v1:0` | Aligned with Anthropic |\n| **6** | **Vertex AI** | Service Account JSON | `gemini-2.5-pro` | Enterprise Gemini via Google Cloud |\n| **7** | **Groq** | `gsk_...` | `deepseek-r1` | Best reasoning model with speed advantage |\n| **8** | **OpenRouter** | `sk-or-...` | `deepseek/deepseek-r1` | Best reasoning model available |\n| **9** | **xAI** | `xai-...` | `grok-code-fast-1` | Latest coding-focused model (Aug 2025) |\n| **10** | **Mistral** | `[a-f0-9]{64}` | `pixtral-large` | Most advanced model (124B params) |\n\n\n\n### **Client-Specific Setup**\n\n#### **Cursor**\n\n1. **Open Cursor Settings:**\n   - Go to `File` → `Preferences` → `Cursor Settings`\n   - Navigate to the `MCP` tab\n   - Click `+ Add` to add a new MCP server\n\n2. **Add MCP Server Configuration:**\n   ```json\n   {\n     \"command\": \"uv\",\n     \"args\": [\"tool\", \"run\", \"mcp-as-a-judge\"],\n     \"env\": {\n       \"LLM_API_KEY\": \"your-openai-api-key-here\",\n       \"LLM_MODEL_NAME\": \"gpt-4.1\"\n     }\n   }\n   ```\n\n   **📝 Configuration Options:**\n   - **LLM_API_KEY**: Required for Cursor (limited MCP sampling)\n   - **LLM_MODEL_NAME**: Optional custom model (see [Supported LLM Providers](#supported-llm-providers) for defaults)\n\n#### **Claude Code**\n\n1. **Add MCP Server via CLI:**\n   ```bash\n   # Set environment variables first (optional model override)\n   export LLM_API_KEY=\"your_api_key_here\"\n   export LLM_MODEL_NAME=\"claude-3-5-haiku\"  # Optional: faster/cheaper model\n\n   # Add MCP server\n   claude mcp add mcp-as-a-judge -- uv tool run mcp-as-a-judge\n   ```\n\n2. **Alternative: Manual Configuration:**\n   - Create or edit `~/.config/claude-code/mcp_servers.json`\n   ```json\n   {\n     \"command\": \"uv\",\n     \"args\": [\"tool\", \"run\", \"mcp-as-a-judge\"],\n     \"env\": {\n       \"LLM_API_KEY\": \"your-anthropic-api-key-here\",\n       \"LLM_MODEL_NAME\": \"claude-3-5-haiku\"\n     }\n   }\n   ```\n\n   **📝 Configuration Options:**\n   - **LLM_API_KEY**: Required for Claude Code (limited MCP sampling)\n   - **LLM_MODEL_NAME**: Optional custom model (see [Supported LLM Providers](#supported-llm-providers) for defaults)\n\n#### **Other MCP Clients**\n\nFor other MCP-compatible clients, use the standard MCP server configuration:\n\n```json\n{\n  \"command\": \"uv\",\n  \"args\": [\"tool\", \"run\", \"mcp-as-a-judge\"],\n  \"env\": {\n    \"LLM_API_KEY\": \"your-openai-api-key-here\",\n    \"LLM_MODEL_NAME\": \"gpt-5\"\n  }\n}\n```\n\n**📝 Configuration Options:**\n- **LLM_API_KEY**: Required for most MCP clients (except GitHub Copilot + VS Code)\n- **LLM_MODEL_NAME**: Optional custom model (see [Supported LLM Providers](#supported-llm-providers) for defaults)\n\n\n\n\n\n## 🔒 **Privacy & Flexible AI Integration**\n\n### **🔑 MCP Sampling (Preferred) + LLM API Key Fallback**\n\n**Primary Mode: MCP Sampling**\n- All judgments are performed using **MCP Sampling** capability\n- No need to configure or pay for external LLM API services\n- Works directly with your MCP-compatible client's existing AI model\n- **Currently supported by:** GitHub Copilot + VS Code\n\n**Fallback Mode: LLM API Key**\n- When MCP sampling is not available, the server can use LLM API keys\n- Supports multiple providers via LiteLLM: OpenAI, Anthropic, Google, Azure, Groq, Mistral, xAI\n- Automatic vendor detection from API key patterns\n- Default model selection per vendor when no model is specified\n\n\n### **🛡️ Your Privacy Matters**\n\n- The server runs **locally** on your machine\n- **No data collection** - your code and conversations stay private\n- **No external API calls when using MCP Sampling**. If you set `LLM_API_KEY` for fallback, the server will call your chosen LLM provider only to perform judgments (plan/code/test) with the evaluation content you provide.\n- Complete control over your development workflow and sensitive information\n\n## 🤝 **Contributing**\n\nWe welcome contributions! Please see [CONTRIBUTING.md](CONTRIBUTING.md) for guidelines.\n\n### **Development Setup**\n\n```bash\n# Clone the repository\ngit clone https://github.com/OtherVibes/mcp-as-a-judge.git\ncd mcp-as-a-judge\n\n# Install dependencies with uv\nuv sync --all-extras --dev\n\n# Install pre-commit hooks\nuv run pre-commit install\n\n# Run tests\nuv run pytest\n\n# Run all checks\nuv run pytest && uv run ruff check && uv run ruff format --check && uv run mypy src\n```\n\n\n## © Concepts and Methodology\n© 2025 OtherVibes and Zvi Fried. The \"MCP as a Judge\" concept, the \"behavioral MCP\" approach, the staged workflow (plan → code → test → completion), tool taxonomy/descriptions, and prompt templates are original work developed in this repository.\n\n\n## Prior Art and Attribution\nWhile “LLM‑as‑a‑judge” is a broadly known idea, this repository defines the original “MCP as a Judge” behavioral MCP pattern by OtherVibes and Zvi Fried. It combines task‑centric workflow enforcement (plan → code → test → completion), explicit LLM‑based validations, and human‑in‑the‑loop elicitation, along with the prompt templates and tool taxonomy provided here. Please attribute as: “OtherVibes – MCP as a Judge (Zvi Fried)”.\n\n## ❓ FAQ\n\n### How is “MCP as a Judge” different from rules/subagents in IDE assistants (GitHub Copilot, Cursor, Claude Code)?\n| Feature | IDE Rules | Subagents | MCP as a Judge |\n|---------|-----------|-----------|----------------|\n| Static behavior guidance | ✓ | ✓ | ✗ |\n| Custom system prompts | ✓ | ✓ | ✓ |\n| Project context integration | ✓ | ✓ | ✓ |\n| Specialized task handling | ✗ | ✓ | ✓ |\n| Active quality gates | ✗ | ✗ | ✓ |\n| Evidence-based validation | ✗ | ✗ | ✓ |\n| Approve/reject with feedback | ✗ | ✗ | ✓ |\n| Workflow enforcement | ✗ | ✗ | ✓ |\n| Cross-assistant compatibility | ✗ | ✗ | ✓ |\n  - References: [GitHub Copilot Custom Instructions](https://docs.github.com/en/copilot/how-tos/configure-custom-instructions/add-repository-instructions), [Cursor Rules](https://docs.cursor.com/en/context/@-symbols/@-cursor-rules), [Claude Code Subagents](https://docs.anthropic.com/en/docs/claude-code/sub-agents)\n\n### How does the Judge workflow relate to the tasklist? Why do we need both?\n- Tasklist = planning/organization: tracks tasks, priorities, and status. It doesn’t guarantee engineering quality or readiness.\n- Judge workflow = quality gates: enforces approvals for plan/design, code diffs, tests, and final completion. It demands real evidence (e.g., unified Git diffs and raw test output) and returns structured approvals and required improvements.\n- Together: Use the tasklist to organize work; use the Judge to decide when each stage is actually ready to proceed. The server also emits next_tool guidance to keep progress moving through the gates.\n\n### If the Judge isn’t used automatically, how do I force it?\n- In your prompt: \"use mcp-as-a-judge\" or \"Evaluate plan/code/test using the MCP server mcp-as-a-judge\".\n- VS Code: Command Palette → \"MCP: List Servers\" → ensure \"mcp-as-a-judge\" is listed and enabled.\n- Ensure the MCP server is running and, in your client, the judge tools are enabled/approved.\n\n### How do I select models for sampling in VS Code?\n- Open Command Palette (Cmd/Ctrl+Shift+P) → \"MCP: List Servers\"\n- Select \"mcp-as-a-judge\" → \"Configure Model Access\"\n- Check your preferred model(s) to enable sampling\n\n\n\n## 📄 **License**\n\nThis project is licensed under the MIT License (see [LICENSE](LICENSE)).\n\n## 🙏 **Acknowledgments**\n\n- [Model Context Protocol](https://modelcontextprotocol.io/) by Anthropic\n- [LiteLLM](https://github.com/BerriAI/litellm) for unified LLM API integration\n\n---\n\n",
  "bytes": 16597,
  "sha": "746b92ff83f078f0a0eea90f76a720e035b18d3b55241ffaab50ed56afd126d6",
  "repo_slug": "othervibes/mcp-as-a-judge",
  "fonte": "repo",
  "truncated": false,
  "api": "https://agentalog.com/api/listings/mcp_io_github_othervibes_mcp_as_a_judge_20338851/readme"
}