Back to the catalog

io.github.LumabyteCo/clarifyprompt

AI prompt optimization for 58+ platforms across 7 categories with custom platforms

Open source Open in the app JSON README (API)

About

AI prompt optimization for 58+ platforms across 7 categories with custom platforms

Details

Kind
MCP servers
Topic
No topic detected
Publisher
lumabyteco
Origin
official
Category
ferramentas
Transport
local
Version
1.1.3
Stars
12
Forks
2
Last push
2026-07-14T17:07:39Z
Repository state
ativo
Language
TypeScript
License
Apache-2.0
Added
2026-08-29 03:02:01
Updated
2026-08-29 03:02:01
Origin id
io.github.LumabyteCo/clarifyprompt

README

# ClarifyPrompt MCP

[![npm version](https://img.shields.io/npm/v/clarifyprompt-mcp.svg)](https://www.npmjs.com/package/clarifyprompt-mcp)
[![ghcr.io](https://img.shields.io/badge/ghcr.io-clarifyprompt--mcp-2496ED?logo=docker&logoColor=white)](https://github.com/LumabyteCo/clarifyprompt-mcp/pkgs/container/clarifyprompt-mcp)
[![evals](https://github.com/LumabyteCo/clarifyprompt-mcp/actions/workflows/evals.yml/badge.svg?branch=main)](https://github.com/LumabyteCo/clarifyprompt-mcp/actions/workflows/evals.yml)
[![License: Apache-2.0](https://img.shields.io/badge/License-Apache%202.0-blue.svg)](https://opensource.org/licenses/Apache-2.0)
[![Node.js](https://img.shields.io/badge/node-%3E%3D18-brightgreen.svg)](https://nodejs.org/)
[![Listed on Glama](https://img.shields.io/badge/Listed%20on-Glama-blueviolet)](https://glama.ai/mcp/servers/LumabyteCo/clarifyprompt-mcp)

A **context-aware MCP prompt compiler** that transforms vague prompts into platform-optimized prompts for 60+ AI platforms across 7 categories — grounded in your workspace signals (CLAUDE.md, AGENTS.md, .cursorrules, package.json), resolved intent, and the capabilities of the target model.

Send a raw prompt. ClarifyPrompt gathers the right context, resolves what you're actually trying to do, and returns a version specifically optimized for Midjourney, DALL-E, Sora, Runway, Higgsfield, ElevenLabs, Claude, ChatGPT, Cursor, or any of the 60+ supported platforms — with the right syntax, parameters, structure, and grounding.

> **New in 1.15.0:** **Nano Banana** (Google Gemini 2.5 Flash Image) is now a **built-in image platform** — `optimize_prompt(platform: "nano-banana")` compiles image prompts in its native style (natural-language scene direction, photographic terms, edit-preserving-identity phrasing, in-image text). Plus **latest-model compatibility across every provider**: `claude-sonnet-5`, `gpt-5`/o-series, and Gemini reject `temperature` and/or `max_tokens`; the client now sends the right parameters (proactively for known reasoning ids, and learns the rest from a `400`). Verified live against Anthropic, OpenAI, Gemini, and Ollama Cloud. See [CHANGELOG.md](./CHANGELOG.md).
>
> **New in 1.14.1:** **Portable-by-default text output** — a `chat`/`document`/`code` prompt with no explicit platform now stays platform-neutral instead of quietly defaulting to Claude's idioms (XML tags); name a platform to opt into vendor-specific tuning. Plus the MCP Apps compose panel now shows a `for <platform>` badge and a clean **Your prompt → Optimized** before/after (with a `show changes` toggle) instead of an always-on diff. See [CHANGELOG.md](./CHANGELOG.md).
>
> **New in 1.14.0:** **An interactive compose panel via MCP Apps.** In hosts that speak the `io.modelcontextprotocol/ui` extension (Claude Desktop, ChatGPT, Cursor, VS Code, …), `compose_prompt` renders a live panel: original-vs-optimized view, all six **critique scores**, the pipeline **stages**, and **Accept / Revise** actions — Accept records the outcome into ClarifyPrompt's memory loop, Revise sends your feedback back into the chat. One self-contained `ui://` resource; hosts without the extension see zero change. See [CHANGELOG.md](./CHANGELOG.md).
>
> **New in 1.13.0:** **Plain-language rewrites.** Optimized prompts now stick to common, everyday words instead of drifting into formal vocabulary ("use", never "utilize") — specificity comes from concrete details, not fancier synonyms. `critique_prompt` gained a 6th default dimension, **`plain_language`**, so `auto_revise` loops correct register drift automatically. Also fixed: an explicit `mode` (e.g. `simple`) is no longer silently dropped for small local models under compact system-prompt shaping. See [CHANGELOG.md](./CHANGELOG.md).

## How It Works

ClarifyPrompt does two things a plain prompt template can't. Every output below is a **real, unedited capture** from `optimize_prompt` run against this repo (see [Provenance](#provenance) at the end of this section).

**1 — It knows each platform.** Same raw prompt, different target, completely different output:

```text
You write:    "a dragon flying over a castle at sunset"

→ Midjourney  A colossal, majestic dragon with shimmering scales soaring over a towering
              medieval stone castle, dramatic sunset sky with vibrant orange and deep purple
              hues, cinematic fantasy concept art, volumetric lighting, highly detailed
              --ar 16:9 --v 6.1 --s 250 --q 2

→ DALL-E      A majestic dragon with glowing crimson scales soars over a towering medieval
              stone castle, silhouetted against a vibrant orange and purple sunset sky.
              Rendered in a high-fantasy digital art style with dramatic, warm lighting and
              highly detailed textures, wide aspect ratio.

→ Nano Banana A majestic dragon with deep crimson scales and a leathery, bat-like wingspan
              glides through the warm, golden-hour sky just above a towering medieval castle
              made of weathered grey stone. ... Frame this as a wide cinematic landscape shot
              using a 24mm lens at f/8 for deep depth of field, camera positioned at a
              slightly elevated three-quarter angle... Aspect ratio 16:9.
```

Midjourney gets `--ar/--v/--s/--q` flags; DALL-E and Nano Banana get flag-free natural language — and Nano Banana layers in photographic direction (lens, f-stop, camera angle) and explicit mood, its documented style. Same idea, each platform's native dialect.

**2 — It knows what you're working on.** This is the part a template can't fake. Drop a vague one-liner while editing `src/transport.ts` *in this very repo*, and the engine grounds it in your real workspace — `package.json`, git state, the active file — and resolves intent **before** it shapes the output:

```text
You write:    "add a configurable request timeout to the http transport"
              · active file: src/transport.ts   · resolved intent: production-code
              · grounded in: active-file · workspace-meta · git-state · environment ·
                target-model · platform-hints

→ Cursor      Implement a configurable request timeout for the HTTP transport in
              `src/transport.ts`.
              Requirements:
              1. Add a new environment variable `CLARIFYPROMPT_HTTP_TIMEOUT` … (default 30000 ms)
              2. Apply this timeout to all incoming requests in the streamable-http transport
              …
              5. Preserve existing behavior for stdio and a2a transports
              …
              The implementation should be added to the streamable-http section of
              `startTransport()`.
              (excerpted — the full rewrite has 7 numbered requirement groups)
```

Nothing in that one-line prompt mentioned the `CLARIFYPROMPT_HTTP_*` naming convention, the `startTransport()` entry point, or the stdio/a2a transports it must preserve — the engine read those from the active file and `package.json` and folded them in. That's the difference between *rephrasing a prompt* and *compiling it against context*.

**3 — It can run the whole pipeline.** clarify → ground/optimize → critique → revise, in one `compose_prompt` call — see **Previously in 1.4.0 — the composable pipeline** below.

> <a name="provenance"></a>**Provenance.** Image outputs captured via `glm-5.2:cloud`, the grounded code output via `qwen3-coder:480b-cloud` — both [Ollama](https://ollama.com) cloud models served over Ollama's OpenAI-compatible endpoint (`LLM_API_URL=http://localhost:11434/v1`), run through `optimize_prompt` against this repo on 2026-06-22 (the Nano Banana capture added 2026-07-03, same `glm-5.2:cloud` model). ClarifyPrompt is model-agnostic (any OpenAI-compatible API, local or hosted); outputs are model-dependent — yours will differ in wording, not in structure.

## What's new in 1.15.0

**Nano Banana, built in.** Google's Gemini 2.5 Flash Image ("Nano Banana") is now a first-class image platform — `optimize_prompt(category: "image", platform: "nano-banana")` compiles your idea into its native prompting style: full-sentence scene direction (not keyword piles), photographic terminology for camera/lens/depth, explicit lighting, edit-phrasing that preserves subject identity, multi-reference character consistency, and reliable in-image text. Like every image platform, ClarifyPrompt compiles the prompt; you send it to the model.

**Latest-model compatibility, every provider.** Thinking-enabled models reject parameters clarifyprompt always sent: `claude-sonnet-5` and OpenAI reasoning models reject `temperature`; `gpt-5` / o-series also reject `max_tokens` (they require `max_completion_tokens`). Every call to them used to fail and degrade to the original prompt. Now the client sends the right body — **proactively** for well-known reasoning ids (no wasted round-trip) and, for anything the hints don't recognize (including future models), it **learns from the `400` and retries**. Models that accept the standard parameters are byte-identical. Verified live against Anthropic (`claude-sonnet-5`), OpenAI (`gpt-5`), Gemini (`gemini-flash-latest`), and Ollama Cloud (`glm-5.2:cloud`). Reasoning models think a lot — bump `LLM_TIMEOUT_MS` (the 30s default is often too short).

## What's new in 1.14.1

**Portable by default.** When you optimize a text prompt (`chat`, `document`, `code`) **without naming a platform**, ClarifyPrompt now returns platform-neutral output — clean, portable structure that works in any assistant — instead of quietly defaulting to Claude's idioms (its `<task>`/`<context>` XML tags). Name a platform (`platform: "claude"`, `"chatgpt"`, … any of the 60) to opt into that platform's specific tuning. Creative categories (image/video/voice/music) are unchanged: their output needs a concrete platform format, so the flagship default (Midjourney, Runway, …) still applies.

**Clearer compose panel.** The MCP Apps panel now shows a `for <platform>` (or `general purpose`) badge, renders your original prompt as a labeled **Your prompt** block above the optimized output, and shows the optimized prompt plainly — with a `show changes` toggle for the word-level diff — instead of an always-on diff.

## What's new in 1.14.0

**`compose_prompt` now has a face.** ClarifyPrompt ships an [MCP Apps](https://github.com/modelcontextprotocol/ext-apps) panel (extension `io.modelcontextprotocol/ui`) that supporting hosts render inline next to the tool result:

- **Original vs optimized, as a word-level diff** — see exactly what the compiler changed.
- **Critique, visualized** — all six dimensions (clarity, specificity, intent_alignment, format_fitness, length_appropriateness, plain_language) as score bars, with the verdict and the per-call `stages` audit trail as badges.
- **Accept** — one click records `save_outcome(accepted)` from the panel, feeding the few-shot memory loop, and quietly tells the model the prompt was accepted.
- **Revise…** — type what should change; the panel sends it back into the chat so the model re-composes.
- **Clarification-aware** — when the pre-clarify stage stops the chain with questions, the panel renders them (with suggested answers) instead of a diff.

Zero-risk rollout: the panel is one self-contained HTML resource (`ui://clarifyprompt/compose-panel.html`, inline CSS/JS — the extension sandbox blocks external requests) linked from `compose_prompt`'s `_meta.ui`. Hosts without the extension ignore it entirely; the text + `structuredContent` output is byte-identical. Runs on the existing SDK ^1.29 floor. New deterministic `npm run test:apps` battery locks the wiring.

Also new: the eval harness gained a **`max_reading_grade`** check — a deterministic Flesch–Kincaid ceiling that locks 1.13.0's plain-language behavior as a measurable gate (formal-register slop scores ~20+; plain rewrites ~3–6).

## What's new in 1.13.0

**Plain-language rewrites, end to end.** LLMs handle common, everyday wording more reliably than formal synonyms of the same meaning — and small local models, ClarifyPrompt's default targets, benefit the most. This release bakes that into every stage that shapes output wording:

- **The optimizer prefers common words.** A new core principle in the shared system prompt ("USE COMMON WORDS") applies to all 7 category strategies and both `optimize_prompt` and `ground_prompt`: never swap in a rarer word where a common one carries the same meaning. Detail means more information, not fancier words — specificity, structure, and constraints are untouched.
- **`critique_prompt` gained a 6th default dimension: `plain_language`.** It penalizes needlessly formal or rare vocabulary where a simpler word would do. Because the rewrite pass applies every suggestion from dimensions scoring below 7, `auto_revise` loops now correct register drift for free. Custom `criteria` overrides are unaffected.
- **Fixed: explicit `mode` no longer silently dropped for small local models.** Compact system-prompt shaping used to trim the mode instructions entirely — so `mode: "simple"` had no effect on 3B-class models. Every mode now survives compact shaping as a one-line rule.
- **Two new eval fixtures** guard the behavior: `31-plain-language-vocabulary` (optimized output must not contain formal-register words) and `32-shape-compact-keeps-mode` (the mode line reaches small models).

## What's new in 1.12.1

The real fix for issue [#3](https://github.com/LumabyteCo/clarifyprompt-mcp/issues/3): **thinking-channel models now reliably produce optimized prompts** instead of intermittently returning empty content. Both `gpt-oss:20b-cloud` and `glm-5.2:cloud` went from empty ~40% of runs to **0%**.

Re-investigating from scratch overturned the documented root cause. It was **never** "Ollama's `/v1` shim drops the harmony final channel." These models spend their `max_tokens` budget on the **thinking channel first** and never reach the final channel — so `content` comes back `""` (worse at higher reasoning effort). Two levers, applied together because different families honor different ones:

- **A `max_tokens` floor (8192) for detected reasoning models** — the *universal* lever. It attacks the root cause directly, so it works regardless of which thinking knob a family respects. It's a ceiling, not a target: short answers finish early, so no added latency.
- **`reasoning_effort: "low"`** — for families that respect it (gpt-oss), also trimming latency/cost. Tune with **`LLM_REASONING_EFFORT`** (`low` | `medium` | `high`).

The levers are genuinely family-specific: **gpt-oss** honors `reasoning_effort` but ignores Ollama's `think`; **glm** is the exact opposite — it ignores `reasoning_effort`, so only the budget floor saves it.

**Detection is robust, not a hardcoded model list** (which would rot as new models ship). "Is this a thinking model?" is answered, cached per model, by: (1) **the runtime itself** — Ollama's `/api/show` reports a `thinking` capability (this is how `minimax-m3:cloud` is detected, with no name match); (2) **response-learning** — any reasoning trace, or empty-content-with-tokens, marks that model thereafter (works for any provider); (3) a small name hint as last resort. Non-reasoning models stay byte-identical, and the name-agnostic empty-content retry is the final backstop. Validated on `gpt-oss:20b-cloud`, `glm-5.2:cloud`, and `minimax-m3:cloud` (all 0% empty on the first call).

> The previously-proposed "switch to Ollama's native `/api/chat`" was a **dead end** — `/api/chat` with `think:false` *still* returns empty content for gpt-oss (it ignores it), and it would have added a fragile second code path.

## What's new in 1.12.0

Step #7 — the final step — of the [MCP modernization roadmap](./docs/audits/mcp-completeness-2026-05.md): ClarifyPrompt now speaks **A2A (Agent-to-Agent)**, so other agents can call it to compile prompts. **stdio stays the default; nothing about existing setups changes.**

Set `CLARIFYPROMPT_TRANSPORT=a2a` and ClarifyPrompt comes up as a discoverable A2A peer on Node's built-in `http` (the only new dependency is the official `@a2a-js/sdk`, which itself pulls just `uuid`):

| Endpoint | Purpose |
|---|---|
| `GET /.well-known/agent-card.json` | **Agent card** — discovery: identity, capabilities, the `compile-prompt-for-platform` skill |
| `POST /a2a` | A2A **JSON-RPC 2.0**: `message/send`, `message/stream` (SSE), `tasks/get`, `tasks/cancel`, … |
| `GET /health` | Liveness probe |

```bash
CLARIFYPROMPT_TRANSPORT=a2a CLARIFYPROMPT_HTTP_PORT=3000 npx clarifyprompt-mcp
# → card:  http://127.0.0.1:3000/.well-known/agent-card.json
# → a2a:   POST http://127.0.0.1:3000/a2a   (message/send · message/stream)
```

The whole roadmap pays off here — one incoming A2A message flows through the same compose pipeline, and the primitives built in earlier steps map straight onto A2A semantics:

- **Compile** — a `message/send` with the raw prompt (plain text, or JSON `{ prompt, platform?, category?, … }`) returns a task whose **artifact** carries the optimized prompt (text) plus the full structured compose result (data).
- **Streaming** (1.10.0 progress → A2A) — `message/stream` emits `status-update` events as each pipeline stage runs, then the artifact, over **SSE**.
- **Cancellation** (1.10.0 AbortSignal → A2A) — `tasks/cancel` aborts the in-flight compose within milliseconds and reports a terminal `canceled` state.
- **Clarification** (1.9.0 elicitation → A2A) — clarify is **off by default** for one-shot peers; opt in with `pre_clarify: 'auto' | 'always'` and an ambiguous prompt pauses the task in A2A's first-class **`input-required`** state with the questions (readable text + structured data). Answer on the same task and it compiles.

Configure the public base URL advertised in the card with `CLARIFYPROMPT_A2A_BASE_URL` (handy behind a proxy); port/host are shared with `streamable-http`. New deterministic `npm run test:a2a` battery drives card discovery, a live compile, the clarify round-trip, and SSE streaming.

## What's new in 1.11.0

Step #6 of the [MCP modernization roadmap](./docs/audits/mcp-completeness-2026-05.md): a pluggable **transport factory** — ClarifyPrompt can now serve over **Streamable HTTP**, the runway toward A2A and remote MCP hosts. **stdio stays the default; nothing about existing setups changes.**

### Transports

Set `CLARIFYPROMPT_TRANSPORT`:

| Value | Behaviour |
|---|---|
| `stdio` (default) | One server over stdin/stdout — exactly as before |
| `streamable-http` | MCP Streamable HTTP over Node's built-in `http` (no new deps): stateful sessions (`mcp-session-id`), SSE streaming, a `/health` probe |
| `a2a` | Serve as an **A2A (Agent-to-Agent)** peer — agent card, JSON-RPC + SSE (see 1.12.0 above) |

HTTP knobs (in `streamable-http` / `a2a` mode): `CLARIFYPROMPT_HTTP_PORT` (3000), `CLARIFYPROMPT_HTTP_HOST` (`127.0.0.1` — localhost-only by default), `CLARIFYPROMPT_HTTP_PATH` (`/mcp`, streamable-http only).

```bash
CLARIFYPROMPT_TRANSPORT=streamable-http CLARIFYPROMPT_HTTP_PORT=3000 npx clarifyprompt-mcp
# → POST http://127.0.0.1:3000/mcp  ·  GET http://127.0.0.1:3000/health
```

Tool/resource registration moved into an exported `createServer()` factory: stdio gets one server, **streamable-http gets one per session** (the SDK-recommended, GHSA-safe pattern — never shares a server across HTTP clients). New deterministic `npm run test:http` battery drives a full HTTP session.

## What's new in 1.10.0

Step #5 of the [MCP modernization roadmap](./docs/audits/mcp-completeness-2026-05.md), stable core: **`compose_prompt` is cancellable and reports live progress**. Model-agnostic, opt-in, fully back-compat.

### Cancellation

An `AbortSignal` is plumbed through the entire LLM path (`simpleGenerate` → `chat` → `fetch`, combined with the per-call timeout) and every engine stage. When a client sends `notifications/cancelled` for a `compose_prompt` call, the in-flight model request aborts immediately and the revise loop stops at the next stage boundary — instead of running every iteration to completion. The signal reaches `fetch` regardless of which model/provider is configured.

### Progress

Include a `progressToken` in the `compose_prompt` request `_meta` and the server emits `notifications/progress` at each stage (clarify / optimize / ground / critique) with a monotonic counter and a human message like `optimizing prompt [iter 2/3]`. Hosts can show a live status on a long multi-iteration compose. No token → no notifications, zero overhead.

### Why not MCP tasks (yet)

Roadmap #5 named the MCP **tasks** API. It's still `experimental/` in the SDK ("may change without notice"), its reference is ~600 lines, and no current client speaks the `tasks/*` protocol — so a full implementation would be unusable off-by-default code today. The real value (cancellable + progress-reporting compose) is delivered here on **stable** primitives; the experimental async-task wrapper is deferred to land with #7 (A2A), which the `AbortSignal` groundwork here already sets up. New deterministic `npm run test:cancel` battery locks the behavior.

## What's new in 1.9.0

Step #4 of the [MCP modernization roadmap](./docs/audits/mcp-completeness-2026-05.md): **`clarify_with_user` can elicit answers through the host's native form UI**. Opt-in, fully back-compat.

### Interactive clarification

Pass `elicit: true`. On a client that supports MCP elicitation, the clarifying questions become a real form:

- each question is a field, `options` become enum dropdowns, and each `suggestedAnswer` is the field default (one-click accept);
- the user answers inline; the engine returns `answers: [{ question, dimension, answer, usedSuggested }]` with `elicited: true`.

Without `elicit`, on a non-capable client, or if the round-trip errors, the tool returns the same raw-questions JSON it always has — every existing caller is unaffected. `decline` / `cancel` are surfaced via `elicitationAction`.

This turns clarification from "here's a JSON blob of questions, you render it" into a first-class interactive moment in hosts like Claude Desktop. The mapping lives in a small pure module (`src/engine/clarification/elicit.ts`), reusable by `compose_prompt`'s pre-clarify stage later. New deterministic `npm run test:elicit` battery (pure helpers + a live mock-client round-trip) locks it.

## What's new in 1.8.0

Step #3 of the [MCP modernization roadmap](./docs/audits/mcp-completeness-2026-05.md): the engine's read surfaces become **browseable resource templates** with **argument autocompletion**. No tool or engine behavior changes.

### Resource templates

Four templates join the static `clarifyprompt://categories`, each backed by an existing engine getter:

| URI template | What it reads |
|---|---|
| `clarifyprompt://platforms/{category}/{id}` | One platform's full config — `resources/list` enumerates all 60+ as individual URIs |
| `clarifyprompt://traces/{date}` | Optimization-trace summary index for a UTC day |
| `clarifyprompt://packs/{id}` | One loaded knowledge pack's metadata |
| `clarifyprompt://memory/facts/{scope}` | Live remembered facts under a scope |

MCP hosts with a resource browser (Claude Desktop, Cursor) now get a navigable tree instead of a single static blob.

### Autocomplete

`completion/complete` resolves the template variables: `{category}` → the 7 category ids, `{id}` → platform ids scoped by the chosen `{category}`, `{date}` → days with traces, pack ids, memory scopes. (MCP completion applies to prompt args + resource-template variables only — not tool inputs; ClarifyPrompt registers no prompts, so it lives on the templates.)

### Capabilities

The server now advertises `resources` (with templates) and `completions` at `initialize`. New deterministic `npm run test:resources` battery locks the surface.

## What's new in 1.7.1

Patch fixing [#3](https://github.com/LumabyteCo/clarifyprompt-mcp/issues/3): a **silent empty optimized prompt** from models whose answer didn't land in `content`.

- **Reads all three thinking-channel field names** (`reasoning` / `thinking` / `reasoning_content`) — fixes DeepSeek / qwen-thinking and similar.
- **Retries once, then fails loudly** when content is empty regardless of any thinking field. This covers the real issue #3 case: gpt-oss harmony output over Ollama's `/v1` shim generates tokens (`completion_tokens > 0`) but returns `content: ""` with no thinking field. The engine now degrades to the original prompt + a surfaced `error` instead of returning blank.
- **Genuinely recovering** gpt-oss harmony output (via Ollama's native `/api/chat`) was tracked as a follow-up — **resolved in 1.12.1**, which proved the `/api/chat` path a dead end and fixed the actual root cause (a `max_tokens` floor + `reasoning_effort` for reasoning models; see the 1.12.1 notes above).
- **New deterministic `npm run test:thinking`** battery locks the regression with mocked responses (no live cloud dependency).

Verified: `test:thinking`, reasoning battery (gpt-oss degrades loudly; the genuine reasoner `kimi-k2-thinking:cloud` still returns real content), integration, day2, evals, wire.

## What's new in 1.7.0

Step #2 of the [MCP modernization roadmap](./docs/audits/mcp-completeness-2026-05.md): the entire tool surface migrated off the deprecated `server.tool()` shorthand (removed in SDK 2.0) onto `server.registerTool()`. **No engine behavior changes; full back-compat.**

### What hosts get

- **Titles** — every tool has a human-readable display name ("Forget a fact", not `memory_forget`).
- **Behavior annotations** — all 23 tools declare `readOnlyHint` / `destructiveHint` / `idempotentHint` / `openWorldHint`. The three destructive tools (`memory_forget`, `unload_pack`, `unregister_platform`) are flagged for confirmation UIs; the seven read-only inspectors are flagged safe-to-call-freely; the seven tools that reach the network (LLM / embeddings / web search) carry `openWorldHint: true`.
- **Structured output** — every tool declares an `outputSchema` and returns `structuredContent` alongside the JSON text. Schemas are permissive by design (all-optional, passthrough) — they document the shape without ever rejecting engine output.

### Back-compat

Text content is byte-identical for every tool — including the three array-returning `list_*` tools, whose text stays a bare array while `structuredContent` wraps it in an object per the MCP spec. Error returns unchanged. Verified: wire 7/7, integration 9/9, day2, 26/27 evals with zero output-validation errors.

### Found during verification

[#3](https://github.com/LumabyteCo/clarifyprompt-mcp/issues/3) — cloud `gpt-oss` thinking-channel responses can yield an empty `optimizedPrompt` (remote API change exposing a pre-existing field-name gap in `client.ts`; fix targeted for 1.7.1).

## What's new in 1.6.8

Housekeeping release closing the loops the 1.6.5→1.6.7 cascade opened. **No engine code, MCP tool surface, platform, or env-var changes.**

### Changed

- **CI matrix now tests Node 24** (current active LTS, EOL Apr 2028) alongside 18/20/22 across Ubuntu + macOS. The matrix previously tested two EOL Node versions but not the current LTS at all. Verified before merge that the native deps (`better-sqlite3` + `sqlite-vec`) load and function on Node 24.16.0 in a toolchain-free `node:24-slim` container. `engines` stays `>=18` — maximum compatibility, and we test what we claim.
- **Publish runner moved Node 20 → 22**, keeping an EOL runtime off the release-critical path (matches the Dockerfile base).

### Process

- **New ship-check `CP-13 — lockfile regeneration safety`** encodes the lesson from the 1.6.5→1.6.6→1.6.7 cascade: a single `npm install --package-lock-only` silently dropped 4 of 5 `sqlite-vec` platform binaries (broke Linux CI) *and* pulled a within-caret `better-sqlite3` bump that dropped Node 20 prebuilds (broke the Docker build). The check mandates full `npm install` on dep changes, a lockfile diff for dropped platform deps + native-dep version jumps, and a local slim-Docker load gate. Dogfooded on this release.

## What's new in 1.6.7

Dockerfile patch. **No engine code, MCP tool surface, platform, or env-var changes.**

### Fixed

- **`CI / docker build` failed on 1.6.6** with `npm error gyp ERR! find Python`. Root cause: `better-sqlite3@12.10.0` ([released 2026-05](https://github.com/WiseLibs/better-sqlite3/releases/tag/v12.10.0)) explicitly removed prebuilt binaries for Node.js v20 and v23 because Node 20 reached EOL in April 2026. The 1.6.6 lockfile regen pulled 12.10.0 within the `^12.9.0` caret, and `node:20-slim` doesn't have Python + a C++ toolchain to compile from source. Bumped the Dockerfile base to `node:22-slim` — current active LTS, still has working prebuilts.
- The non-Docker CI build matrix (Node 18 / 20 / 22 across macOS + Ubuntu) still passes because regular runners can compile-from-source as fallback. Only the slim Docker image stumbles.

### Verified locally

`docker build` → green. Container can `require('better-sqlite3')` + `require('sqlite-vec')` cleanly. All 5 `sqlite-vec` platform binaries still in `package-lock.json` (1.6.6's fix held).

## What's new in 1.6.6

Lockfile + harness patch following 1.6.5. **No engine code, MCP tool surface, platform, or env-var changes.** Ships the MCP-completeness audit doc.

### Fixed

- **`package-lock.json` lost 4 of 5 `sqlite-vec` platform binaries during the 1.6.5 SDK bump.** My local `npm install --package-lock-only` retained only the maintainer's `sqlite-vec-darwin-arm64` binary. `npm ci` on CI's Ubuntu runners failed with `no such module: vec0` because `sqlite-vec-linux-x64` wasn't in the lock. End-user `npm install clarifyprompt-mcp@1.6.5` was **never affected** (the npm tarball doesn't ship a lockfile; users resolve platforms at install time). Regenerated with full `npm install` so all 5 platforms (`darwin-arm64`, `darwin-x64`, `linux-arm64`, `linux-x64`, `windows-x64`) are back.
- **Eval harness HTML report writer crashed on ERRORED entries** (`evals/run.mjs:729`). The pre-existing renderer assumed every non-skipped, non-filtered run had an `evaluation.checks` field, but errored runs carry an `error` field instead. Added an explicit errored-status branch — the harness now degrades gracefully and exits cleanly even when fixtures error.

### Bundled docs

- **[`docs/audits/mcp-completeness-2026-05.md`](./docs/audits/mcp-completeness-2026-05.md)** — diagnostic audit of the engine's MCP surface against the current SDK + spec. Tool-by-tool registration table, resource gap analysis, SDK feature delta (1.12 → 1.29 → 2.0-alpha), capability declarations, transport refactor sketch, A2A feasibility note, and a sequenced 7-step modernization roadmap. The artifact behind next-session planning. No engine changes prescribed inline.

### Numbers

- 5 sqlite-vec platforms in lockfile (was 1). `npm audit --production`: 0 vulnerabilities (unchanged). Tools: 23 (unchanged). Eval fixtures: 30 (unchanged).

## What's new in 1.6.5

Security patch. **No engine code changes, no MCP tool surface changes, no platform changes, no env-var changes.**

### Fixed

- **CVE-2026-0621** — ReDoS in `@modelcontextprotocol/sdk`'s `UriTemplate` regex (patched in SDK `1.25.2`). The previous `^1.12.1` floor *allowed* vulnerable resolutions on stale npm caches; bumped to `^1.29.0` so the floor itself is patched.
- **GHSA-345p-7cg4-v4c7** — Shared server/transport instances leak cross-client response data (patched in SDK `1.26.0`). Not exploitable in practice for ClarifyPrompt (one host = one server instance) but the vulnerable code is now out of the dependency graph entirely.
- **7 transitive vulnerabilities** (2 moderate, 5 high) in the SDK's bundled HTTP-transport substack (`hono`, `express-rate-limit`, `fast-uri`, `ip-address`, `path-to-regexp`, `qs`, `@hono/node-server`). Cleared via `npm audit fix`. Never affected runtime — ClarifyPrompt is stdio-only and doesn't load the HTTP transport — but they were noise in users' `npm audit` reports and made the install look unsafe.

### Numbers

- `npm audit --production` → **0 vulnerabilities** (was 2 SDK CVEs + 7 transitive).
- `package-lock.json`: **net −336 lines** (the old caret was pulling in heavy unused HTTP-transport ancillaries; the fix swapped them for slimmer alternates).
- Tools: 23 (unchanged). Platforms: 60+ (unchanged). Eval fixtures: 30 (unchanged).
- Wire test + integration battery + day2 + reasoning + 29/30 evals pass against the new floor on local Ollama. The one eval fail (`analyzer-creative-media`) is a pre-existing qwen-coder-7b classifier flake — verified SDK-independent by stash-reverting and re-running.

### Why the floor bump matters

`^1.12.1` was misleading documentation — caret resolution was actually pulling SDK `1.27.1` for any fresh `npm install` since early 2026. The floor bump aligns the declared baseline with what `npm` was already doing for most users while guaranteeing the floor for users on stale caches. It also positions us for the eventual `2.0.0-alpha` migration when that line stabilizes (the modern SDK deprecates `.tool()` / `.prompt()` / `.resource()` shorthand registration in favor of `registerTool()` / `registerPrompt()` / `registerResource()` with title + outputSchema + annotations).

## What's new in 1.6.4

Docs + process patch. **No engine, MCP tool, or platform changes** — but a meaningful cleanup of the pack-distribution model.

### Pack registry consolidated back into the engine repo

`LumabyteCo/clarifyprompt-packs` (the separate community-pack registry created in `1.3` with the right principle but at the wrong scale) has been **archived**. Its three starter packs already lived in this repo's [`packs/`](./packs/) folder; the registry was meant to be the canonical home but in practice everything always shipped from here via the npm tarball. The drift caught up: `higgsfield-creative-handbook` shipped in `1.6.2` and never made it to the registry, even though the registry's own README told users to fetch packs from there.

Net result of 1.6.4:

- **Single source of truth.** `packs/*.md` knowledge packs + `packs/platforms/*.yaml` platform configs all live in `clarifyprompt-mcp` and ship in the npm tarball.
- **New top-level [Knowledge packs](#knowledge-packs) section** in this README explains the loading model (`load_knowledge_pack({source: "<url-or-path>", scope: ...})`), the three starter packs + Higgsfield, the scope semantics, and how to contribute.
- **New [`packs/README.md`](./packs/README.md)** — pack authoring guide (frontmatter schema, chunk boundaries, quality bar). Lifted from the archived registry so the content isn't lost.
- **Tombstone redirect on the archived repo.** Anyone visiting `clarifyprompt-packs` lands on a banner pointing here.

### When does the split come back?

When there's a forcing function: a community PR queue on packs alone, pack count >20, or divergent licensing/governance. Until then the maintenance cost of keeping two repos in sync wasn't paying for an audience that hadn't materialized.

### Numbers

- **Tools:** 23 (unchanged).
- **Platforms:** 60+ (unchanged).
- **Bundled knowledge packs:** 4 (`anthropic-brand-voice`, `higgsfield-creative-handbook`, `nextjs-14-best-practices`, `sox-compliance`) — same as 1.6.2/1.6.3, just newly canonical.
- **Eval fixtures:** 30 (unchanged).
- **Tarball size:** unchanged from 1.6.3.

## What's new in 1.6.3

Patch. The 1.6.2 CI tag-push run surfaced two real issues — fixed here without changing any engine code.

### Fixed

- **`evals/fixtures/28-context-includes-git-state.yaml`** previously asserted `git_branch_present: true`, but GitHub Actions checks out in detached-HEAD mode where `bundle.git.branch` is correctly `undefined` (only the SHA + recent commits are populated). Relaxed to assert `bundle_has_git: true` only — that's what's actually invariant across local + CI environments.
- **`evals/fixtures/17-critique-strong-prompt-accepts.yaml`** asserted `verdict: accept` + `overall_score_min: 7` on a strong prompt. gpt-4o-mini's judge calibrates stricter than qwen2.5-coder:7b's, and occasionally returned a malformed `overall` field that the parser defaulted to 0 → verdict=reject. The fixture's *real intent* is to verify engine wiring (5+ dimensions, the standard dimension names present, no harness error) — not to compare judge calibration across models. Dropped the verdict + tight score assertions; kept the wiring-level checks.
- **README Glama badge** swapped from inline `<img>` (sometimes broken via GitHub's camo proxy) to a shields.io text-link badge that's stable across all rendering surfaces.

### Notes

- **No engine code changes.** No new MCP tools (still 23). No platform changes (still 60+). No env-var changes.
- **Eval baselines unchanged on local Ollama.** This is a CI-specific hardening — local runs against qwen-coder-7b produced the same results before and after.
- **The CI publish-gate failure that appeared on the v1.6.2 tag push** was downstream of the eval failure (`Wait for evals workflow` step blocked publish). Now that the underlying fixtures don't false-fail on gpt-4o-mini + detached-HEAD CI, the publish gate clears too.

## What's new in 1.6.2

Patch. Two additive ships, both no-code-changes from the engine's perspective:

### Higgsfield creative-handbook knowledge pack

`packs/higgsfield-creative-handbook.md` — a community-style markdown pack documenting Higgsfield's actual conventions: model-selection rules (which of the 13 models for which use case), Soul ID character-training workflow, camera-move vocabulary, prompt-structure pattern (long-form prose, not keyword tags), multi-reference editing, Marketing Studio modes, common pitfalls (don't translate Midjourney flags verbatim), output specs.

Load it explicitly:

```
load_knowledge_pack source="https://raw.githubusercontent.com/LumabyteCo/clarifyprompt-mcp/main/packs/higgsfield-creative-handbook.md"
```

…or, since it ships in the npm tarball, point at the installed copy. The Context Curator grounds Higgsfield-targeted prompts in this pack's chunks automatically via semantic retrieval. See the [Knowledge packs](#knowledge-packs) section for the full loading + scoping model.

### `npm run matrix` — multi-model eval matrix runner

`evals/matrix.mjs` runs `npm run eval` sequentially against N models and stitches the results into one side-by-side HTML (`evals/matrix.html` by default). Lights up the model-class-gated fixtures (`shape-small-local-model` / `shape-mid-tier-model` / `shape-reasoning-model`) that single-model runs skip, and exposes deltas like "qwen-7b fails analyzer-creative-media but gpt-4o-mini passes it" in a glance.

```bash
npm run matrix -- --models qwen2.5-coder:7b-instruct-q4_K_M,gpt-oss:20b-cloud,glm-5.2:cloud
```

Outputs a dark-themed table — rows = fixtures, columns = models, cells = pass / fail / skip / errored with tooltips showing which checks failed.

Companion fix: `evals/run.mjs` gains a `--json-out <path>` flag that writes structured per-model results (matrix.mjs uses it; CI agents can use it too).

### Numbers

- **No tool surface change.** Still 23 MCP tools.
- **No platform count change.** 60+ platforms (`packs/platforms/*.yaml` unchanged).
- **30 → 30 fixtures** (no new fixtures; matrix is tooling, not coverage).
- **Tarball grows ~10 KB** for the knowledge pack. `evals/matrix.mjs` is NOT in the tarball — it's a maintainer/contributor tool, not a runtime artifact.

## What's new in 1.6.1

Patch release. Adds **[Higgsfield](https://higgsfield.ai)** as a target platform in both `image` and `video` categories. No code changes — pure YAML platform-pack additions and one eval fixture.

Higgsfield is a multi-model creative platform that exposes its own MCP server at `https://mcp.higgsfield.ai/mcp`. Inside one connection you get:

- **Image**: Soul 2.0, Soul Cinema, Soul Cast (character-consistent), Flux 2, Seedream 5, Nano Banana Pro, GPT Image 2
- **Video**: Cinema Studio, Sora 2, Veo 3.1, Kling 3.0, WAN 2.6, Seedance 2.0
- **Workflows**: Soul ID character training, Lipsync Studio, UGC Factory, Marketing Studio, virality_predictor

The 1.6.1 ClarifyPrompt platform entries surface Higgsfield's model identifiers and prompt-style conventions (long-form natural-language prose; composition + lighting + textures + mood; up to 4K images / 15 s video / Soul ID for character consistency) as syntax hints to the curator.

**Recommended pattern:** install both `clarifyprompt-mcp` AND Higgsfield's MCP in your client (Claude Desktop / Cursor / AI Butler / Claude Code). Use `optimize_prompt(platform: 'higgsfield', ...)` or `compose_prompt(platform: 'higgsfield', ...)` to compile, then pass the compiled prompt to Higgsfield's `generate_image` / `generate_video` tool. MCPs compose at the client; ClarifyPrompt stays at the "compile" layer.

29 → 30 eval fixtures. Same MCP tool surface as 1.6.0 (23 tools, 1 resource). No env-var changes.

## What's new in 1.6.0

Four targeted additions across the engine's four pillars (memory / agentic / models / context), each shipped behind real eval fixtures. **3 new MCP tools** (23 total). Fully back-compat with 1.5.x — no removed tools, no removed fields, no required env-var changes.

### Memory — explicit fact CRUD (`memory_remember`, `memory_forget`, `memory_list_facts`)

Before 1.6, facts only entered persistent memory via *reflection on `save_outcome`* — implicit, LLM-extracted, after-the-fact. 1.6 adds the explicit path:

- **`memory_remember`** — directly insert a `(subject, predicate, object)` triple with explicit confidence. Source tagged `user:explicit`. Auto-embedded for future semantic retrieval.
- **`memory_forget`** — soft-delete (bi-temporal `invalidated_at`) a fact by id. Idempotent: re-forgetting an already-invalidated fact is a no-op and returns `success: false` cleanly.
- **`memory_list_facts`** — list live facts in a scope (default `user`), optionally filtered by predicate. Sorted by most-recently-observed.

This closes the obvious UX gap where the engine could only learn from *outcomes* — now users can say "remember I prefer X" directly.

### Agentic — `compose_prompt`'s new `max_iterations` revise loop

`compose_prompt` used to revise *once* (the critique's `improvedPrompt` replaced the optimization, if the verdict wasn't `accept`). 1.6 adds a loop:

```json
{ "prompt": "...", "post_critique": true, "auto_revise": true, "max_iterations": 3 }
```

Each iteration after the first re-runs `optimize` + `critique` on the previous iteration's improved prompt. Stops at `verdict=accept`, no improvedPrompt to feed back, or the cap. `pre_clarify` only runs once (no point re-asking on a rewrite). The response includes a new `iterations` field showing how many fired. Hard cap of 5 to prevent cost runaways.

### Models — per-stage model routing

Each compose stage can now target a different model:

```json
{
  "prompt": "...",
  "clarify_model": "qwen2.5-coder:7b-instruct-q4_K_M",
  "optimize_model": "claude-sonnet-5",
  "critique_model": "gpt-4o-mini"
}
```

Run clarify on a cheap local model, optimize on the big-budget frontier model, critique on the cheap judge. The override flows through every layer — `optimization.metadata.model` and `critique.judgeModel` in the response reflect the actual model that ran each stage.

### Context — git-state + environment signals

Two new signal collectors feed the Context Curator:

- **`bundle.git`** — current branch, short SHA, dirty flag, last 5 commit titles. Lets the engine ground prompts in "what you're iterating on" without you spelling it out. Detected via `git rev-parse` / `git status` / `git log`; fails soft when cwd isn't a repo.
- **`bundle.environment`** — `nowIso` / `weekday` / `timezone` (IANA from `Intl.DateTimeFormat`). Helps with time-sensitive prompts ("send this email tomorrow"). Pure JS, never fails.

Both are low-utility candidates in the curator (won't dominate budget) but surface as grounding sources when relevant.

### Eval coverage

23 → **29 fixtures** (6 new):
- 24 `memory-remember-persists` / 25 `memory-forget-invalidates` — Me1 CRUD round-trip
- 26 `compose-loop-iterates` — A1 loop infrastructure (new `iterations_min` / `iterations_max` checks)
- 27 `compose-per-stage-models-honored` — M1 per-stage routing (new `optimization_model_eq` / `critique_model_eq` checks)
- 28 `context-includes-git-state` / 29 `context-includes-environment-time` — C1 + C4 signals (new `bundle_has_git` / `bundle_has_environment` / `git_branch_present` checks)

Local baseline on `qwen2.5-coder:7b`: **25 passed / 1 failed / 3 skipped / 97% avg**. The lone failure remains the persistent `analyzer-creative-media` model-class signal (untouched).

## What's new in 1.5.2

The first release where CI's eval gate (against `gpt-4o-mini`) drove the diff. Three real fixes that the gate caught the moment we wired in the `OPENAI_API_KEY` secret:

- **Memory store now supports any embedding dimension** ([#2](https://github.com/LumabyteCo/clarifyprompt-mcp/issues/2)). The persistent vec table was hardcoded to 768 dims (the nomic-embed-text default), so anyone configuring `EMBED_MODEL=text-embedding-3-small` (1536), `voyage-3` (1024), `embed-english-v3.0` (1024), or any non-768 model would hit `Dimension mismatch: expected 768, got N` on the first `memory_search` call. The store now derives the table name from the embedder's actual dimension and creates the dim-specific table at boot. Existing 768-dim installs are unaffected.
- **`LLM_TIMEOUT_MS` env-var override** on the LLM client. Default stays at 30s; users on slow hosted models can bump it. The eval workflow uses 120s for `gpt-4o-mini`.
- **Eval harness hardened** — no longer crashes when a tool throws an exception (the SDK returns plain-text error responses; the harness used to `JSON.parse` them and die). One bad fixture no longer tanks the whole run.
- **Live evals badge.** The `evals.yml` workflow runs on every push to main. The `[![evals]](...)` badge at the top of this README is its real-time status. Currently green at 20/0/3 · 100% on `gpt-4o-mini`.

No new MCP tools. No env-var surface changes (only an added optional `LLM_TIMEOUT_MS`). Fully back-compat with 1.5.x.

## What's new in 1.5.1

A patch release on top of 1.5.0. Pure docs + ship-process improvements; **runtime behavior is identical to 1.5.0**.

- **README marketing surfaces refreshed** — the 1.5.0 release shipped with the README still on 1.4.0 in three places (headline blockquote, "What's new in X" heading, "cumulative through X" annotation). Every other version surface (`package.json`, `package-lock.json`, `server.json`, `src/index.ts`, `CHANGELOG`) was correct, but the prose drifted because nothing automated touched it. 1.5.1 fixes that.
- **Two new ship-check audits** — `CP-11` (README marketing-surface coherence) hard-fails if any of the three above don't reference the current `package.json#version`. `CP-12` (Platform-pack format validity) parses every `packs/platforms/*.yaml` and asserts schema validity. CP-11 was promoted to the user-scoped (cross-project) ship-check skill the same day, so future projects benefit too.
- **No code changes.** No new MCP tools. No new env vars. Same tarball anatomy as 1.5.0 plus a few hundred bytes of CHANGELOG.

## What's new in 1.5.0

**Built-in platforms become declarative.** The 58+ hardcoded TypeScript platform arrays move to `packs/platforms/*.yaml` — adding a built-in platform is now a YAML edit, not a TS edit. The TypeScript layer becomes a runtime loader with a hardcoded fallback table. Malformed YAML can never soft-brick the server.

```
packs/platforms/
  chat.yaml       9 platforms
  code.yaml       9
  document.yaml   8
  image.yaml     10
  music.yaml      4
  video.yaml     11
  voice.yaml      7
  README.md      contributor docs
```

To add a new built-in platform: append an entry to the relevant category file, run `npm run build`, open a PR. No TS edit required. Custom-platform-via-runtime (`register_platform`) still works identically for user-installed platforms.

- **Memory-layer eval coverage.** The eval harness now supports `setup: [{tool, args}, ...]` — a list of MCP tool calls executed BEFORE the main `input`. Two new fixtures use it: one loads a knowledge pack inline and verifies the chunk surfaces in `grounding.sources` after the embed → store → retrieve → curate → ground pipeline; the other proves vector-search ranking quality. **23 fixtures total** (was 20 in 1.4.0).
- **Test infrastructure modernization.** The integration + Day-2 test batteries used to assert literal version strings (`1.3.0`, `16 tools`) and broke on every bump. Now they read `EXPECTED_VERSION` from `package.json` and assert presence of a tool *set* rather than a tool *count*. Future bumps don't break the tests.
- **Adoption materials.** `docs/adoption/` ships with copy/paste-ready Show HN body, Reddit posts, Twitter thread, awesome-mcp-servers PR template, and catalog submission specs (mcp.so, Smithery, mcp-get, PulseMCP, modelcontextprotocol/servers).
- **One new runtime dep:** `js-yaml` promoted from devDependency for the platform loader (~200 KB).
- **Same MCP tool surface as 1.4.** 20 tools, 1 resource. No new tools; no removed tools; result shapes unchanged.

## Previously in 1.4.0 — the composable pipeline

Four core operations as first-class MCP tools that compose. Use any tool standalone, or run the whole chain in one call:

```
  ┌─────────────┐     ┌─────────────────────┐     ┌──────────────┐
  │  clarify    │ →   │  ground OR optimize │ →   │   critique   │
  │  (optional) │     │       (core)        │     │  (optional)  │
  └─────────────┘     └─────────────────────┘     └──────────────┘

  one call = compose_prompt(prompt, [sources], post_critique, auto_revise, ...)
```

- **`clarify_with_user`** — Given an ambiguous draft, returns 1–3 targeted clarifying questions, each with a `suggested_answer` you can accept verbatim, optional 2–4 quick-pick `options`, and a `dimension` tag (audience/scope/format/length/tone/constraints/goal/platform). Short-circuits with `clarificationNeeded: false` on confident, well-formed prompts so it pipelines cleanly in front of `optimize_prompt` without a per-call latency tax.
- **`ground_prompt`** — The strict, retrieval-augmented variant of `optimize_prompt`. Caller-provided sources are pinned at the **highest** priority — above project rules, above pinned instructions — and tracked individually in the trace as `user-source:N`. Strict mode: zero non-empty sources → error, no silent fall-through. Per-source body cap (4000 chars) so a single huge paste can't dominate the budget.
- **`critique_prompt`** — LLM-as-judge. Scores a candidate prompt 0–10 across 5 default dimensions (clarity, specificity, intent_alignment, format_fitness, length_appropriateness) — or your own criteria — with per-dimension rationale + concrete suggestions, an overall score, and a verdict (`accept` / `revise` / `reject`). Below `revise_threshold` (default 7.0) it also returns an `improvedPrompt` you can drop in. Use it pre-flight ("is this prompt good enough for the expensive model?"), postmortem ("was the prompt the cause?"), or to A/B-pick the best of N optimization variants.
- **`compose_prompt`** — One MCP call runs the canonical pipeline. Auto-decides the ground vs. optimize branch from whether you passed `sources`. `pre_clarify: 'auto' | 'always' | 'never'`. `post_critique: true` adds a judge pass. `auto_revise: true` replaces `final_prompt` with the rewrite when the verdict isn't `accept`. Returns a per-stage `stages` audit array so the caller sees exactly what ran.
- **Eval harness v0** — Deterministic regression tests under `evals/`. 20 YAML fixtures cover analyzer, shape, intent-overlay, grounding, clarify, critique, ground, and compose surfaces. `npm run eval` produces a console summary + self-contained dark-themed HTML report. Multi-model matrix is just bash: run `LLM_MODEL=... npm run eval -- --report-path evals/report-X.html` per model.
- **CI-gated evals (opt-in)** — When `OPENAI_API_KEY` is set as a repo secret, the eval harness runs in CI against `gpt-4o-mini` as a release gate. Off by default; nothing leaves your machine without the secret.
- **5 new MCP tools** (20 total). `optimize_prompt` also gains a `userProvidedSources` injection point — both `ground_prompt` and `compose_prompt` use it under the hood, but it's available directly if you want explicit control without the strict-mode validation.

> Carried over from 1.3: persistent memory + knowledge packs + reflective learning. The curator continues to score and fit grounding sources into the target model's remaining window. `explain_last_curation` still gives you a per-call breakdown of selected vs. rejected candidates with reasons.

## What's in the box (cumulative through 1.15.0)

- **Context Engine** — auto-gathers workspace rules (`CLAUDE.md`, `AGENTS.md`, `.cursorrules`, `.clinerules`, `clarify.md`), detects frameworks and languages from `package.json` and sibling manifests, tracks an active file excerpt, and maintains a per-session ring buffer of recent optimizations **and their outcomes**.
- **Unified `PromptAnalyzer`** — one LLM call produces `{ category, intent, recommendedMode, confidence }` together. 10 intents: `production-code`, `brand-voice`, `stakeholder-comm`, `data-extract`, `creative-media`, `technical-spec`, `analysis`, `quick-draft`, `exploration`, `unknown`. Intent beats surface keywords on ambiguity.
- **Target-model-aware prompt shaping** — system prompt, `maxTokens`, and `temperature` adapt to the downstream LLM's context window and the resolved intent. Small local models get a compact prompt; Claude/GPT-4/Gemini get the full richness.
- **Grounding Context (single, priority-ordered)** — user pinned instructions → project rules → active file → prior accepted examples → web search → workspace metadata → target-model hints → custom platform instructions → built-in syntax hints. No more parallel context silos.
- **Session retrieval (save_outcome)** — the caller reports `accepted | edited | rejected` per optimization; similar accepted outputs in the same session get injected as few-shot examples into future similar prompts. Backed by persistent memory (SQLite + sqlite-vec), so accepted outcomes survive restarts.
- **Local JSONL tracing** — every optimization writes a structured trace line (now with `shape`, `groundingSources`, `error` fields) to `$CLARIFYPROMPT_HOME/traces/YYYY-MM-DD.jsonl`. **Nothing is uploaded.** Toggle via `CLARIFYPROMPT_TRACE=off`.
- **Unified `$CLARIFYPROMPT_HOME`** — one env var for everything ClarifyPrompt writes. Legacy `CLARIFYPROMPT_CONFIG_DIR` / `CLARIFYPROMPT_DATA_DIR` still work (deprecation hint, silenceable).
- **Three transports** — `stdio` (default), `streamable-http` (MCP over Node `http`, stateful sessions + `/health`), and `a2a` (an **Agent-to-Agent** peer: agent card, JSON-RPC `message/send` + SSE `message/stream`, task cancellation, `input-required` clarification). One `CLARIFYPROMPT_TRANSPORT` env var; stdio behavior is byte-identical to before.
- **60+ platforms, 7 categories, custom platforms** — the original core is unchanged and fully backward-compatible.
- **Any LLM, any provider.** One code path works with **any OpenAI-compatible API** — Ollama (local + cloud), LM Studio, vLLM, OpenAI, Google Gemini, xAI Grok, Groq, Mistral, DeepSeek, Cohere, Perplexity, Together, Fireworks, OpenRouter — plus **Anthropic Claude** directly. Reasoning models (`o1/o3/o4`, `deepseek-reasoner`, `gpt-oss`, `*-thinking`) are auto-detected and given a larger token budget so they actually produce content. [See 15+ pre-configured provider examples below](#provider-examples).
- **Apache-2.0, forever.** Open-source core, no relicensing.

## Quick Start

### With Docker

Pull the published image from GitHub Container Registry (multi-arch: `amd64` + `arm64`, with signed provenance + SBOM):

```bash
docker pull ghcr.io/lumabyteco/clarifyprompt-mcp:latest
```

All config is passed at **run time** — nothing is baked into the image, so the image is safe to share and contains no secrets:

```bash
# stdio (for MCP hosts that launch the container)
docker run --rm -i \
  -e LLM_API_URL=http://host.docker.internal:11434/v1 \
  -e LLM_MODEL=qwen2.5:7b \
  -e CLARIFYPROMPT_HOME=/data \
  -v clarifyprompt-data:/data \
  ghcr.io/lumabyteco/clarifyprompt-mcp:latest

# or serve over HTTP / A2A
docker run --rm -p 3000:3000 \
  -e CLARIFYPROMPT_TRANSPORT=a2a -e CLARIFYPROMPT_HTTP_HOST=0.0.0.0 \
  -e LLM_API_URL=http://host.docker.internal:11434/v1 -e LLM_MODEL=qwen2.5:7b \
  ghcr.io/lumabyteco/clarifyprompt-mcp:latest
```

> Mount a volume at `CLARIFYPROMPT_HOME` to persist memory, traces, and packs across runs. Pass `LLM_API_KEY` / `EMBED_API_KEY` as `-e` env vars (or `--env-file`) at run time — never bake them into an image.

### With Claude Desktop

Add to your `claude_desktop_config.json`:

```json
{
  "mcpServers": {
    "clarifyprompt": {
      "command": "npx",
      "args": ["-y", "clarifyprompt-mcp"],
      "env": {
        "LLM_API_URL": "http://localhost:11434/v1",
        "LLM_MODEL": "qwen2.5:7b"
      }
    }
  }
}
```

### With Claude Code

```bash
claude mcp add clarifyprompt -- npx -y clarifyprompt-mcp
```

Set the environment variables in your shell before launching:

```bash
export LLM_API_URL=http://localhost:11434/v1
export LLM_MODEL=qwen2.5:7b
```

### With Cursor

Add to your `.cursor/mcp.json`:

```json
{
  "mcpServers": {
    "clarifyprompt": {
      "command": "npx",
      "args": ["-y", "clarifyprompt-mcp"],
      "env": {
        "LLM_API_URL": "http://localhost:11434/v1",
        "LLM_MODEL": "qwen2.5:7b"
      }
    }
  }
}
```

### With AI Butler

[AI Butler](https://github.com/LumabyteCo/aibutler) is a self-hosted
personal AI agent runtime — single Go binary, multi-channel chat, MCP
ecosystem hub. Drop ClarifyPrompt into its `mcp.servers` config and
the agent picks up all 23 tools as native capabilities, callable from
any channel (web chat, terminal, Telegram, Slack, etc.). AI Butler
discovers tools dynamically via MCP's `tools/list`, so adding /
removing tools in ClarifyPrompt updates the agent's surface
automatically — no config edits needed on the butler side.

Edit `~/.aibutler/config.yaml`:

```yaml
configurations:
  mcp:
    servers:
      - name: clarifyprompt
        command: clarifyprompt-mcp
        env:
          LLM_API_URL: "http://localhost:11434/v1"
          LLM_MODEL: "qwen3-vl:8b"
```

Restart AI Butler. The boot log confirms the tools are wired in:

![AI Butler boot log: mcp: connected to clarifyprompt, 1/1 servers connected, then "Ready. Press Ctrl+C to stop." Verified live integration.](docs/screenshots/aibutler/01-boot-log.png)

The agent enumerates the full surface on request — every tool prefixed
with `clarifyprompt.`:

![AI Butler webchat showing the agent listing the clarifyprompt tools with one-line descriptions](docs/screenshots/aibutler/02-tool-listing.png)

> 📸 **Screenshots above are from a 1.2-era integration (11 tools).**
> Current `1.6.x` exposes **23 tools** — `optimize_prompt`,
> `clarify_with_user`, `ground_prompt`, `critique_prompt`,
> `compose_prompt`, plus the management / inspection / memory
> tools (`memory_search`, `memory_remember`, `memory_forget`,
> `memory_list_facts`, knowledge-pack tools, traces, custom
> platforms, etc.). AI Butler picks them up automatically via the
> MCP `tools/list` discovery; no config changes needed.

#### Drive the Context Engine end-to-end

You can preview what the engine *would* gather (without running the
optimization) using `inspect_context`:

![Context Engine preview — analyzer output (Category=code, Intent=production-code, Recommended Mode=detailed, Confidence=Medium), session history, and the priority-ordered grounding stack the engine would merge into the system prompt. Closing takeaway about a language mismatch the engine detected between workspace (JS) and prompt (TypeScript).](docs/screenshots/aibutler/03-inspect-context.p

More