{
  "markdown": "<p align=\"center\">\n  <img src=\"https://raw.githubusercontent.com/ellmos-ai/open-compute-mcp/main/assets/wappen.jpg\" alt=\"open-compute MCP server emblem\" width=\"400\">\n</p>\n\n# open-compute-mcp\n\n**npm launcher for the [open-compute](https://github.com/ellmos-ai/open-compute) MCP server** —\nmodel-agnostic **computer-use** tools exposed over the Model Context Protocol (MCP).\n\n**EN** | [DE](README_de.md)\n\n[![npm version](https://img.shields.io/npm/v/open-compute-mcp.svg)](https://www.npmjs.com/package/open-compute-mcp)\n[![npm downloads](https://img.shields.io/npm/dt/open-compute-mcp.svg)](https://www.npmjs.com/package/open-compute-mcp)\n[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT)\n[![Node.js](https://img.shields.io/badge/node-%3E%3D18-brightgreen.svg)](https://nodejs.org/)\n[![Node.js CI](https://img.shields.io/badge/tests-21%20passed-brightgreen.svg)](https://github.com/ellmos-ai/open-compute-mcp/actions)\n[![MCP Enabled](https://img.shields.io/badge/MCP-server-blue.svg)](https://modelcontextprotocol.io)\n[![Platform](https://img.shields.io/badge/platform-Windows%20%7C%20Linux%20%7C%20macOS-lightgrey.svg)](https://github.com/ellmos-ai/open-compute-mcp)\n[![Privacy: Zero-Egress](https://img.shields.io/badge/privacy-100%25%20Offline%20%7C%20Zero--Egress-blue.svg)](SECURITY.md)\n[![Security: Safety-Gated](https://img.shields.io/badge/security-Operator%20Ceiling%20%7C%20Safety--Gated-green.svg)](SECURITY.md)\n[![Ecosystem: ellmos-ai](https://img.shields.io/badge/ecosystem-ellmos--ai-blueviolet.svg)](https://github.com/ellmos-ai)\n[![Umbrella: open-bricks](https://img.shields.io/badge/umbrella-open--bricks-indigo.svg)](https://github.com/open-bricks)\n[![LLM Ready](https://img.shields.io/badge/LLM-ready-success.svg)](https://github.com/ellmos-ai/open-compute-mcp/blob/main/llms.txt)\n\n📦 **[View on npm →](https://www.npmjs.com/package/open-compute-mcp)** • 📋 **[Security Policy](SECURITY.md)** • ⚖️ **[Licenses](THIRD_PARTY_LICENSES.md)** • 🤖 **[LLM Context (llms.txt)](llms.txt)**\n\n<a href=\"https://glama.ai/mcp/servers/ellmos-ai/open-compute-mcp\"><img src=\"https://raw.githubusercontent.com/ellmos-ai/open-compute-mcp/main/assets/glama-badge.jpg\" alt=\"Glama: open-compute-mcp — A license, B maintenance\" width=\"100%\"></a>\n\n---\n\n### Quick Navigation\n\n- [✨ Key Capabilities](#key-capabilities)\n- [🏗️ Architecture](#architecture)\n- [🛠️ Tools (16)](#tools)\n- [🚀 Use with an MCP Client](#use-with-an-mcp-client)\n- [🔄 Safe Interaction & Signal Lifecycle](#safe-interaction--signal-lifecycle)\n- [⚙️ Configuration](#configuration-environment-variables)\n- [🔒 Safety & Security](#safety)\n- [🌐 ellmos-ai Ecosystem](#ellmos-ai-ecosystem)\n\n---\n\n> [!NOTE]\n> **AI Assistant / Agent Integration**: This repository contains an [`llms.txt`](llms.txt) file providing structured, machine-readable specifications of tools, safety modes (`OC_SAFETY_MODE`), and client configuration examples for RAG crawlers and autonomous agent frameworks.\n\nThe MCP **client is the reasoner** (no API key, model-agnostic): it calls `capture`\nto see the screen, then acts with `do` / `click_name` / `invoke`. This is the keyless\nMode-A loop of open-compute, but as native tool-calls.\n\n## Key Capabilities\n\n1. **State-bound Perception & Window Targeting:** Captures/trees return one-shot observation IDs; window enumeration returns stable window/process IDs and issued tokens. WGC remains the GPU-window fallback.\n2. **Fail-closed Action Execution:** Coordinates consume one observation and exact window binding; UIA names resolve exact-first; text is segmented with focus checks and character-count postconditions.\n3. **Leased Signal Overlay & Abort Control:** The glowing border/cursor signal has owner/session metadata, a bounded TTL, turn-end cleanup, and immediate human abort.\n4. **Multimodal Collaboration & Voice Notes:** Push-to-talk voice recording (`talk`), screen chat messaging (`chat`), directory monitoring (`watch_dir`), and macro replay (`rec_replay`).\n\n## Architecture\n\n```mermaid\ngraph TD\n    A[\"AI Reasoner<br/>(Claude / Antigravity / Cursor)\"] -- \"MCP stdio (JSON-RPC)\" --> B[\"npx open-compute-mcp<br/>(Node.js Launcher)\"]\n    B -- \"Spawns via uvx\" --> C[\"open-compute Python Engine<br/>(GitHub @ main)\"]\n    C -- \"Screenshots / WGC\" --> D[\"Windows Display\"]\n    C -- \"UIA / Mouse / Keys\" --> E[\"Windows Desktop Apps\"]\n    C -- \"Glowing Border & Cursor\" --> F[\"Signal Overlay UI\"]\n\n    subgraph Safety Gate\n        C -. \"OC_SAFETY_MODE<br/>(confirm / read_only / allow_all)\" .-> C\n        C -. \"OC_DENY<br/>(hard action blacklist)\" .-> C\n    end\n```\n\n> This package is a **thin launcher**. It contains no server logic — it spawns the\n> **Python** open-compute server (pulled from GitHub) and pipes MCP stdio through.\n> Real screen capture and input require the **interactive Windows desktop session**.\n\n## Requirements\n\n- **Python 3.10+** and **[uv](https://docs.astral.sh/uv/)** on the host. The default\n  launch uses `uvx` to fetch open-compute (with the `mcp` extra) **from GitHub** on\n  first run — the `mcp` extra tracks the GitHub repo, so this works regardless of\n  PyPI release timing.\n- **Windows** for real capture/input (mss + UIA). Other platforms import the tools\n  but cannot drive a desktop.\n\n## Tools\n\n| Tool | Purpose |\n|---|---|\n| `capture` | Return one-shot observation metadata plus an image (optionally one exact window). |\n| `do` | Execute a safety-gated action; coordinates require `observation_id` + issued window descriptor/token. |\n| `tree` | Return UIA elements and a one-shot observation ID for their coordinates. |\n| `click_name` | Exact-first, ambiguity-safe click in a required issued window, with score/alternatives. |\n| `invoke` | Exact-first, click-free UIA activation in a required issued window. |\n| `list_windows` | List stable window/process IDs, exact titles, issued tokens, rects and centers. |\n| `get_screen_size` | Virtual-desktop geometry + per-monitor breakdown (read-only). |\n| `watch_dir` | Watch directories for file-system changes. |\n| `push_status` | Feed-manager status (read-only). |\n| `rec_replay` | Replay a `.clirec` macro (needs the optional `clirec` package). |\n| `signal_show` | Show a configurable pre-action color/text countdown, then the mode-colored overlay, with owner/session lease and bounded TTL. |\n| `signal_hide` | Hide the signal overlay. |\n| `signal_status` | Owner/session/mode/visible/expires_at + pending abort message. |\n| `signal_abort` | Ask the human for a short abort reason; the message is returned for the model. |\n| `chat` | Human→model message about screen content, optionally with screenshot. |\n| `talk` | Push-to-talk voice note → WAV path (hold key, speak, release; STT/TTS model-side). |\n\nAll coordinates are **normalized 0..1** relative to the virtual desktop. Tool\ndescriptions are localized in six languages (`de/en/es/ja/ru/zh`) via `OC_LANGUAGE`.\n\n`do` also accepts the **hold primitives** `mouse_down` / `mouse_up` / `key_down` /\n`key_up` for press-and-hold sequences (rubber-band selection, modifier-held\nclicking, game input); anything still held is released when the server stops.\n`capture(window=...)` falls back to Windows.Graphics.Capture when a plain grab of\na hardware-composited window (Roblox Studio, Blender, a GPU-accelerated browser)\ncomes back all-black — install the `wgc` extra for that.\n\n## Safe Interaction & Signal Lifecycle\n\nThe v0.8 Python engine enforces observe → one action → automatic refresh. Keep\nthe full descriptor or `window_token` from `list_windows`, then pass it as\n`expected_window` together with the latest `observation_id` from `capture` or\n`tree`. `click_name`/`invoke` require that issued window too. Reuse, changed\nstate, focus mismatch, covered windows, and ambiguous UIA targets are rejected\nbefore input. `type` returns requested/sent character\ncounts and complete/partial status without echoing the text. Signals have a\nhard TTL and are removed at action turn end unless `keep_signal=true`.\n\nAn explicit `signal_show` starts the engine's configured pre-action grace\nperiod. The static grace color is distinct from the mode color and the visible\ntext counts down `Start in N Sekunden` once per second. At zero, both phase and\ncolor switch once to active. `signal_status` exposes the same phase, remaining\nseconds, current color, and screenreader label. Duration, grace color, and text\ntemplate come from `OC_SIGNAL_GRACE_SECONDS` / `OC_SIGNAL_CONFIG`; `0` skips the\ncountdown. The design uses no flashing, pulsing, or motion animation.\n\n```mermaid\nsequenceDiagram\n    autonumber\n    actor Reasoner as AI Reasoner (Claude / AGY)\n    participant Launcher as Node.js Launcher (open-compute-mcp)\n    participant Engine as Python Engine (open-compute)\n    participant UI as Windows Desktop / UIA\n    actor Operator as Human Operator\n\n    Note over Reasoner,Operator: Phase 1: Visual Perception & State Inspection\n    Reasoner->>Launcher: capture(window?) / tree()\n    Launcher->>Engine: Forward stdio JSON-RPC\n    Engine->>UI: Grab Screen (mss/WGC) or Read UIA Tree\n    UI-->>Engine: Frame Image / Semantic Element Tree\n    Engine-->>Launcher: Observation ID + normalized response/image\n    Launcher-->>Reasoner: State-bound visual observation\n\n    Note over Reasoner,Operator: Phase 2: Signal Overlay Activation\n    Reasoner->>Launcher: signal_show(mode=\"control\")\n    Launcher->>Engine: Invoke Signal Overlay\n    Engine->>UI: Render static grace color + Start in N seconds\n    UI-->>Operator: Text countdown + accessible window name\n    Engine->>UI: At zero, switch once to the mode color\n\n    Note over Reasoner,Operator: Phase 3: Action Request & Safety Gate\n    Reasoner->>Launcher: do(one action, window token, observation_id) / click_name(target)\n    Launcher->>Engine: Process Action Payload\n    alt OC_SAFETY_MODE == \"confirm\" (Default)\n        Engine-->>Launcher: Status \"needs_confirmation\" (Report Only)\n        Launcher-->>Reasoner: Human confirmation needed\n    else OC_SAFETY_MODE == \"allow_all\" (Isolated VM)\n        Engine->>UI: Execute Mouse/Keyboard / Hold Primitives\n        UI-->>Engine: Action Completed\n        Engine-->>Launcher: Post-observation + window/modal/text postconditions\n        Launcher-->>Reasoner: Action completed; old observation invalid\n    end\n\n    Note over Reasoner,Operator: Phase 4: Emergency Abort or Completion\n    opt Operator Triggers Emergency Abort\n        Operator->>Engine: Hotkey Pressed (Abort Signal)\n        Engine->>UI: Auto-release all held keys/mouse buttons\n        Engine-->>Reasoner: signal_abort message returned\n    end\n    Engine->>UI: Remove overlay on turn end/error/abort (unless keep_signal=true)\n```\n\n## Use with an MCP client\n\n**Via this npm launcher (npx):**\n\n```json\n{\n  \"mcpServers\": {\n    \"open-compute\": {\n      \"command\": \"npx\",\n      \"args\": [\"-y\", \"open-compute-mcp\"]\n    }\n  }\n}\n```\n\n**Directly via Python (uvx), no npm:**\n\n```json\n{\n  \"mcpServers\": {\n    \"open-compute\": {\n      \"command\": \"uvx\",\n      \"args\": [\"--from\", \"open-compute[mcp,local,uia] @ git+https://github.com/ellmos-ai/open-compute.git\", \"open-compute-mcp\"]\n    }\n  }\n}\n```\n\n## Configuration (environment variables)\n\n| Variable | Effect |\n|---|---|\n| `OPEN_COMPUTE_PYTHON` | Path to a `python.exe`; the launcher runs `-m open_compute.mcp_server` with it (use this if you installed open-compute into a specific environment). |\n| `OPEN_COMPUTE_MCP_CMD` | Full command override (whitespace-split), e.g. `python -m open_compute.mcp_server`. |\n| `OPEN_COMPUTE_GIT_REF` | Git ref (branch/tag/sha) to pin for the uvx launch (default: the repo's default branch). |\n| `OPEN_COMPUTE_EXTRAS` | Extras for the default `uvx` launch (default `mcp,local,uia`). |\n| `OC_LANGUAGE` | Language of the tool descriptions: `de`/`en`/`es`/`ja`/`ru`/`zh`. |\n| `OC_SAFETY_MODE` | `confirm` (default) · `read_only` · `allow_all`. |\n| `OC_DENY` | Comma-separated action types always denied (e.g. `type,launch_app`). |\n| `OC_CAPTURE_SCALE` | Resize factor for every capture, `0.05`–`1.0`. **This launcher defaults to `0.5`** (see below); set `1.0` for full resolution. |\n| `OC_CAPTURE_MAX_DIM` | Cap the longest edge in pixels (default off). Setting it suppresses the scale default, so the two never shrink twice. |\n| `OC_CAPTURE_GRAYSCALE` | `1` drops colour. Shrinks the payload, **not** the token count — that follows pixel count alone. |\n| `OC_SIGNAL_TTL` | Hard overlay lease limit in seconds (default 120). |\n| `OC_SIGNAL_IDLE_HIDE` | Additional idle timeout for explicitly kept auto-signals (default 60). |\n| `OC_SIGNAL_GRACE_SECONDS` | Pre-action countdown duration (default 20; `0` starts immediately). |\n| `OC_SIGNAL_CONFIG` | Signal JSON containing `pre_action_grace_color`, `pre_action_grace_label`, and per-mode colors. |\n\n### Capture size — why this launcher halves it by default\n\nA vision model is billed per pixel, and every frame **stays in the conversation**, so a\nfull-HD grab is charged again on each following request. The cost of a session therefore\ngrows with the *square* of the number of screenshots, not linearly.\n\nBecause open-compute's coordinates are **normalized 0..1**, shrinking the image costs\nnothing in click accuracy — `do` works in fractions of the image either way. Only\nlegibility drops, and at `0.5` buttons and field borders stay clearly identifiable; small\nbody text is what gets hard to read.\n\n| Setting | 1920×1080 grab | Cost |\n|---|---|---|\n| `OC_CAPTURE_SCALE=1.0` | full resolution | ~1600 tokens |\n| `OC_CAPTURE_SCALE=0.5` *(this launcher's default)* | 960×540 | ~690 tokens |\n| `OC_CAPTURE_MAX_DIM=768` | 768×432 | ~440 tokens |\n\nThe Python library itself defaults to full resolution — its callers are not necessarily\npaying per pixel. Only this launcher, which exists to serve agents, opts into the smaller\nframe and prints a one-line notice when it does.\n\n**What saves more than any scale factor:** prefer `tree` where\nthe accessibility model carries the content — note that in browsers it usually exposes only\nthe browser chrome, not the page; and use `capture(window=…)` rather than the full desktop.\nCoordinate actions deliberately follow observe → one action → automatic refresh;\ndo not batch multiple coordinate steps against one stale frame.\n\n## Safety\n\nComputer-use is powerful. `OC_SAFETY_MODE` is an operator **ceiling** (`confirm`\ndefault · `read_only` · `allow_all`); a per-call `mode` can only *tighten* it, never\nloosen it. Because MCP stdio has no server→client confirm callback, `confirm` /\n`read_only` **report** an action without performing it. For interactive use, run in\nan **isolated VM/session**, set `OC_SAFETY_MODE=allow_all`, and let your client's\ntool-approval dialog be the human-in-the-loop. `OC_DENY` (comma-separated action\ntypes) is a hard deny list. Treat on-screen content as untrusted (prompt-injection\nrisk).\n\n**Troubleshooting: `do`/`click_name` only ever return `needs_confirmation` and never\nact.** That is the `confirm` ceiling working as designed under stdio MCP. Fix for\ninteractive use: set `\"env\": {\"OC_SAFETY_MODE\": \"allow_all\"}` in the server\nregistration and let the client's tool-approval dialog gate each action (do **not**\nauto-allow `do`/`click_name`/`invoke` there). The env change only takes effect when\nthe server process (re)starts — an already-connected client keeps the old ceiling\nuntil it reconnects.\n\n## License\n\nMIT — see [LICENSE](LICENSE). Part of the open-compute project.\n\n---\n\n## ellmos-ai Ecosystem\n\nThis MCP server is part of the **[ellmos-ai](https://github.com/ellmos-ai)** ecosystem — AI infrastructure, MCP servers, and intelligent tools.\n\n### MCP Server Family\n\n| Server | Tools | Focus | npm |\n|--------|-------|-------|-----|\n| [FileCommander](https://github.com/ellmos-ai/ellmos-filecommander-mcp) | 46 | Filesystem, process management, interactive sessions, cloud-lock-safe operations | [`ellmos-filecommander-mcp`](https://www.npmjs.com/package/ellmos-filecommander-mcp) |\n| [CodeCommander](https://github.com/ellmos-ai/ellmos-codecommander-mcp) | 22 | Code analysis, JSON repair, imports, diffs, regex | [`ellmos-codecommander-mcp`](https://www.npmjs.com/package/ellmos-codecommander-mcp) |\n| [Clatcher](https://github.com/ellmos-ai/ellmos-clatcher-mcp) | 12 | File repair, format conversion, batch operations | [`ellmos-clatcher-mcp`](https://www.npmjs.com/package/ellmos-clatcher-mcp) |\n| [n8n Manager](https://github.com/ellmos-ai/n8n-manager-mcp) | 18 | n8n workflow management via AI assistants | [`n8n-manager-mcp`](https://www.npmjs.com/package/n8n-manager-mcp) |\n| [ControlCenter](https://github.com/ellmos-ai/ellmos-controlcenter-mcp) | 20 | MCP stack discovery, profile management, control plane | [`ellmos-controlcenter-mcp`](https://www.npmjs.com/package/ellmos-controlcenter-mcp) |\n| [Homebase](https://github.com/ellmos-ai/ellmos-homebase-mcp) | 45 | Local-first LLM memory, knowledge, state, routing, swarm orchestration | [`ellmos-homebase-mcp`](https://www.npmjs.com/package/ellmos-homebase-mcp) (alpha) |\n| [ServerCommander](https://github.com/ellmos-ai/ellmos-servercommander-mcp) | 8 | Server operations: health checks, log analysis, deploy dry-runs, mail diagnostics | [`ellmos-servercommander-mcp`](https://www.npmjs.com/package/ellmos-servercommander-mcp) (alpha) |\n| [Blender Use](https://github.com/ellmos-ai/ellmos-blender-use-mcp) | 3 | Headless Blender asset QA and FBX reimport verification | [`ellmos-blender-use-mcp`](https://www.npmjs.com/package/ellmos-blender-use-mcp) (alpha) |\n| **[Open Compute](https://github.com/ellmos-ai/open-compute-mcp)** | **16** | **Model-agnostic computer use: capture, safety-gated actions, Windows UIA, signal overlay & voice/chat** | **[`open-compute-mcp`](https://www.npmjs.com/package/open-compute-mcp)** (alpha) |\n\n### AI Infrastructure & Sibling Tooling\n\n| Project | Description |\n|---|---|\n| [BACH](https://github.com/ellmos-ai/bach) | Local-first text-based OS for LLM agents — 113+ handlers, 550+ tools, SQLite memory |\n| [open-compute](https://github.com/ellmos-ai/open-compute) | Model-agnostic computer-use core powering Open Compute MCP |\n| [clutch](https://github.com/ellmos-ai/clutch) | Provider-neutral LLM orchestration with auto-routing and budget tracking |\n| [rinnsal](https://github.com/ellmos-ai/rinnsal) | Lightweight agent memory, connectors, and automation infrastructure |\n| [ellmos-stack](https://github.com/ellmos-ai/ellmos-stack) | Self-hosted AI research stack (Ollama + n8n + Rinnsal + KnowledgeDigest) |\n| [MarbleRun](https://github.com/ellmos-ai/MarbleRun) | Autonomous agent chain framework for Claude Code |\n| [gardener](https://github.com/ellmos-ai/gardener) | Minimalist database-driven LLM OS prototype (4 functions, 1 table) |\n| [ellmos-tests](https://github.com/ellmos-ai/ellmos-tests) | Testing framework for LLM operating systems (7 dimensions) |\n| [sqlite-transit-sync](https://github.com/ellmos-ai/sqlite-transit-sync) | Safe, redacted, HMAC-verified SQLite snapshot synchronizer |\n| [policy-registry](https://github.com/ellmos-ai/policy-registry) | Hierarchical policy & delegation authority engine |\n\n### Open Bricks Umbrella\n\nOur partner organization **[open-bricks](https://github.com/open-bricks)** bundles AI-native desktop applications — a modern, open-source software suite built for the age of AI. Sibling suites include [DevCenter](https://github.com/dev-bricks/DevCenter), [CodeBox](https://github.com/dev-bricks/CodeBox), [MethodenAnalyser](https://github.com/dev-bricks/MethodenAnalyser), [CleanMarkdown](https://github.com/doc-bricks/CleanMarkdown), and [PDFtoPDFocr](https://github.com/doc-bricks/PDFtoPDFocr).\n\n",
  "bytes": 19578,
  "sha": "5800cddb93c8e1f83c276a2d76ac00e252348675ec5632ca92878b08955dad77",
  "repo_slug": "ellmos-ai/open-compute-mcp",
  "fonte": "repo",
  "truncated": false,
  "api": "https://agentalog.com/api/listings/mcp_io_github_lukisch_open_compute_mcp_a7a45a5b/readme"
}