io.github.LumabyteCo/clarifyprompt
AI prompt optimization for 58+ platforms across 7 categories with custom platforms
Open source Open in the app JSON README (API)
About
AI prompt optimization for 58+ platforms across 7 categories with custom platforms
Details
- Kind
- MCP servers
- Topic
- No topic detected
- Publisher
- lumabyteco
- Origin
- official
- Category
- ferramentas
- Transport
- local
- Version
- 1.1.3
- Stars
- 12
- Forks
- 2
- Last push
- 2026-07-14T17:07:39Z
- Repository state
- ativo
- Language
- TypeScript
- License
- Apache-2.0
- Added
- 2026-08-29 03:02:01
- Updated
- 2026-08-29 03:02:01
- Origin id
io.github.LumabyteCo/clarifyprompt
README
# ClarifyPrompt MCP
[](https://www.npmjs.com/package/clarifyprompt-mcp)
[](https://github.com/LumabyteCo/clarifyprompt-mcp/pkgs/container/clarifyprompt-mcp)
[](https://github.com/LumabyteCo/clarifyprompt-mcp/actions/workflows/evals.yml)
[](https://opensource.org/licenses/Apache-2.0)
[](https://nodejs.org/)
[](https://glama.ai/mcp/servers/LumabyteCo/clarifyprompt-mcp)
A **context-aware MCP prompt compiler** that transforms vague prompts into platform-optimized prompts for 60+ AI platforms across 7 categories — grounded in your workspace signals (CLAUDE.md, AGENTS.md, .cursorrules, package.json), resolved intent, and the capabilities of the target model.
Send a raw prompt. ClarifyPrompt gathers the right context, resolves what you're actually trying to do, and returns a version specifically optimized for Midjourney, DALL-E, Sora, Runway, Higgsfield, ElevenLabs, Claude, ChatGPT, Cursor, or any of the 60+ supported platforms — with the right syntax, parameters, structure, and grounding.
> **New in 1.15.0:** **Nano Banana** (Google Gemini 2.5 Flash Image) is now a **built-in image platform** — `optimize_prompt(platform: "nano-banana")` compiles image prompts in its native style (natural-language scene direction, photographic terms, edit-preserving-identity phrasing, in-image text). Plus **latest-model compatibility across every provider**: `claude-sonnet-5`, `gpt-5`/o-series, and Gemini reject `temperature` and/or `max_tokens`; the client now sends the right parameters (proactively for known reasoning ids, and learns the rest from a `400`). Verified live against Anthropic, OpenAI, Gemini, and Ollama Cloud. See [CHANGELOG.md](./CHANGELOG.md).
>
> **New in 1.14.1:** **Portable-by-default text output** — a `chat`/`document`/`code` prompt with no explicit platform now stays platform-neutral instead of quietly defaulting to Claude's idioms (XML tags); name a platform to opt into vendor-specific tuning. Plus the MCP Apps compose panel now shows a `for <platform>` badge and a clean **Your prompt → Optimized** before/after (with a `show changes` toggle) instead of an always-on diff. See [CHANGELOG.md](./CHANGELOG.md).
>
> **New in 1.14.0:** **An interactive compose panel via MCP Apps.** In hosts that speak the `io.modelcontextprotocol/ui` extension (Claude Desktop, ChatGPT, Cursor, VS Code, …), `compose_prompt` renders a live panel: original-vs-optimized view, all six **critique scores**, the pipeline **stages**, and **Accept / Revise** actions — Accept records the outcome into ClarifyPrompt's memory loop, Revise sends your feedback back into the chat. One self-contained `ui://` resource; hosts without the extension see zero change. See [CHANGELOG.md](./CHANGELOG.md).
>
> **New in 1.13.0:** **Plain-language rewrites.** Optimized prompts now stick to common, everyday words instead of drifting into formal vocabulary ("use", never "utilize") — specificity comes from concrete details, not fancier synonyms. `critique_prompt` gained a 6th default dimension, **`plain_language`**, so `auto_revise` loops correct register drift automatically. Also fixed: an explicit `mode` (e.g. `simple`) is no longer silently dropped for small local models under compact system-prompt shaping. See [CHANGELOG.md](./CHANGELOG.md).
## How It Works
ClarifyPrompt does two things a plain prompt template can't. Every output below is a **real, unedited capture** from `optimize_prompt` run against this repo (see [Provenance](#provenance) at the end of this section).
**1 — It knows each platform.** Same raw prompt, different target, completely different output:
```text
You write: "a dragon flying over a castle at sunset"
→ Midjourney A colossal, majestic dragon with shimmering scales soaring over a towering
medieval stone castle, dramatic sunset sky with vibrant orange and deep purple
hues, cinematic fantasy concept art, volumetric lighting, highly detailed
--ar 16:9 --v 6.1 --s 250 --q 2
→ DALL-E A majestic dragon with glowing crimson scales soars over a towering medieval
stone castle, silhouetted against a vibrant orange and purple sunset sky.
Rendered in a high-fantasy digital art style with dramatic, warm lighting and
highly detailed textures, wide aspect ratio.
→ Nano Banana A majestic dragon with deep crimson scales and a leathery, bat-like wingspan
glides through the warm, golden-hour sky just above a towering medieval castle
made of weathered grey stone. ... Frame this as a wide cinematic landscape shot
using a 24mm lens at f/8 for deep depth of field, camera positioned at a
slightly elevated three-quarter angle... Aspect ratio 16:9.
```
Midjourney gets `--ar/--v/--s/--q` flags; DALL-E and Nano Banana get flag-free natural language — and Nano Banana layers in photographic direction (lens, f-stop, camera angle) and explicit mood, its documented style. Same idea, each platform's native dialect.
**2 — It knows what you're working on.** This is the part a template can't fake. Drop a vague one-liner while editing `src/transport.ts` *in this very repo*, and the engine grounds it in your real workspace — `package.json`, git state, the active file — and resolves intent **before** it shapes the output:
```text
You write: "add a configurable request timeout to the http transport"
· active file: src/transport.ts · resolved intent: production-code
· grounded in: active-file · workspace-meta · git-state · environment ·
target-model · platform-hints
→ Cursor Implement a configurable request timeout for the HTTP transport in
`src/transport.ts`.
Requirements:
1. Add a new environment variable `CLARIFYPROMPT_HTTP_TIMEOUT` … (default 30000 ms)
2. Apply this timeout to all incoming requests in the streamable-http transport
…
5. Preserve existing behavior for stdio and a2a transports
…
The implementation should be added to the streamable-http section of
`startTransport()`.
(excerpted — the full rewrite has 7 numbered requirement groups)
```
Nothing in that one-line prompt mentioned the `CLARIFYPROMPT_HTTP_*` naming convention, the `startTransport()` entry point, or the stdio/a2a transports it must preserve — the engine read those from the active file and `package.json` and folded them in. That's the difference between *rephrasing a prompt* and *compiling it against context*.
**3 — It can run the whole pipeline.** clarify → ground/optimize → critique → revise, in one `compose_prompt` call — see **Previously in 1.4.0 — the composable pipeline** below.
> <a name="provenance"></a>**Provenance.** Image outputs captured via `glm-5.2:cloud`, the grounded code output via `qwen3-coder:480b-cloud` — both [Ollama](https://ollama.com) cloud models served over Ollama's OpenAI-compatible endpoint (`LLM_API_URL=http://localhost:11434/v1`), run through `optimize_prompt` against this repo on 2026-06-22 (the Nano Banana capture added 2026-07-03, same `glm-5.2:cloud` model). ClarifyPrompt is model-agnostic (any OpenAI-compatible API, local or hosted); outputs are model-dependent — yours will differ in wording, not in structure.
## What's new in 1.15.0
**Nano Banana, built in.** Google's Gemini 2.5 Flash Image ("Nano Banana") is now a first-class image platform — `optimize_prompt(category: "image", platform: "nano-banana")` compiles your idea into its native prompting style: full-sentence scene direction (not keyword piles), photographic terminology for camera/lens/depth, explicit lighting, edit-phrasing that preserves subject identity, multi-reference character consistency, and reliable in-image text. Like every image platform, ClarifyPrompt compiles the prompt; you send it to the model.
**Latest-model compatibility, every provider.** Thinking-enabled models reject parameters clarifyprompt always sent: `claude-sonnet-5` and OpenAI reasoning models reject `temperature`; `gpt-5` / o-series also reject `max_tokens` (they require `max_completion_tokens`). Every call to them used to fail and degrade to the original prompt. Now the client sends the right body — **proactively** for well-known reasoning ids (no wasted round-trip) and, for anything the hints don't recognize (including future models), it **learns from the `400` and retries**. Models that accept the standard parameters are byte-identical. Verified live against Anthropic (`claude-sonnet-5`), OpenAI (`gpt-5`), Gemini (`gemini-flash-latest`), and Ollama Cloud (`glm-5.2:cloud`). Reasoning models think a lot — bump `LLM_TIMEOUT_MS` (the 30s default is often too short).
## What's new in 1.14.1
**Portable by default.** When you optimize a text prompt (`chat`, `document`, `code`) **without naming a platform**, ClarifyPrompt now returns platform-neutral output — clean, portable structure that works in any assistant — instead of quietly defaulting to Claude's idioms (its `<task>`/`<context>` XML tags). Name a platform (`platform: "claude"`, `"chatgpt"`, … any of the 60) to opt into that platform's specific tuning. Creative categories (image/video/voice/music) are unchanged: their output needs a concrete platform format, so the flagship default (Midjourney, Runway, …) still applies.
**Clearer compose panel.** The MCP Apps panel now shows a `for <platform>` (or `general purpose`) badge, renders your original prompt as a labeled **Your prompt** block above the optimized output, and shows the optimized prompt plainly — with a `show changes` toggle for the word-level diff — instead of an always-on diff.
## What's new in 1.14.0
**`compose_prompt` now has a face.** ClarifyPrompt ships an [MCP Apps](https://github.com/modelcontextprotocol/ext-apps) panel (extension `io.modelcontextprotocol/ui`) that supporting hosts render inline next to the tool result:
- **Original vs optimized, as a word-level diff** — see exactly what the compiler changed.
- **Critique, visualized** — all six dimensions (clarity, specificity, intent_alignment, format_fitness, length_appropriateness, plain_language) as score bars, with the verdict and the per-call `stages` audit trail as badges.
- **Accept** — one click records `save_outcome(accepted)` from the panel, feeding the few-shot memory loop, and quietly tells the model the prompt was accepted.
- **Revise…** — type what should change; the panel sends it back into the chat so the model re-composes.
- **Clarification-aware** — when the pre-clarify stage stops the chain with questions, the panel renders them (with suggested answers) instead of a diff.
Zero-risk rollout: the panel is one self-contained HTML resource (`ui://clarifyprompt/compose-panel.html`, inline CSS/JS — the extension sandbox blocks external requests) linked from `compose_prompt`'s `_meta.ui`. Hosts without the extension ignore it entirely; the text + `structuredContent` output is byte-identical. Runs on the existing SDK ^1.29 floor. New deterministic `npm run test:apps` battery locks the wiring.
Also new: the eval harness gained a **`max_reading_grade`** check — a deterministic Flesch–Kincaid ceiling that locks 1.13.0's plain-language behavior as a measurable gate (formal-register slop scores ~20+; plain rewrites ~3–6).
## What's new in 1.13.0
**Plain-language rewrites, end to end.** LLMs handle common, everyday wording more reliably than formal synonyms of the same meaning — and small local models, ClarifyPrompt's default targets, benefit the most. This release bakes that into every stage that shapes output wording:
- **The optimizer prefers common words.** A new core principle in the shared system prompt ("USE COMMON WORDS") applies to all 7 category strategies and both `optimize_prompt` and `ground_prompt`: never swap in a rarer word where a common one carries the same meaning. Detail means more information, not fancier words — specificity, structure, and constraints are untouched.
- **`critique_prompt` gained a 6th default dimension: `plain_language`.** It penalizes needlessly formal or rare vocabulary where a simpler word would do. Because the rewrite pass applies every suggestion from dimensions scoring below 7, `auto_revise` loops now correct register drift for free. Custom `criteria` overrides are unaffected.
- **Fixed: explicit `mode` no longer silently dropped for small local models.** Compact system-prompt shaping used to trim the mode instructions entirely — so `mode: "simple"` had no effect on 3B-class models. Every mode now survives compact shaping as a one-line rule.
- **Two new eval fixtures** guard the behavior: `31-plain-language-vocabulary` (optimized output must not contain formal-register words) and `32-shape-compact-keeps-mode` (the mode line reaches small models).
## What's new in 1.12.1
The real fix for issue [#3](https://github.com/LumabyteCo/clarifyprompt-mcp/issues/3): **thinking-channel models now reliably produce optimized prompts** instead of intermittently returning empty content. Both `gpt-oss:20b-cloud` and `glm-5.2:cloud` went from empty ~40% of runs to **0%**.
Re-investigating from scratch overturned the documented root cause. It was **never** "Ollama's `/v1` shim drops the harmony final channel." These models spend their `max_tokens` budget on the **thinking channel first** and never reach the final channel — so `content` comes back `""` (worse at higher reasoning effort). Two levers, applied together because different families honor different ones:
- **A `max_tokens` floor (8192) for detected reasoning models** — the *universal* lever. It attacks the root cause directly, so it works regardless of which thinking knob a family respects. It's a ceiling, not a target: short answers finish early, so no added latency.
- **`reasoning_effort: "low"`** — for families that respect it (gpt-oss), also trimming latency/cost. Tune with **`LLM_REASONING_EFFORT`** (`low` | `medium` | `high`).
The levers are genuinely family-specific: **gpt-oss** honors `reasoning_effort` but ignores Ollama's `think`; **glm** is the exact opposite — it ignores `reasoning_effort`, so only the budget floor saves it.
**Detection is robust, not a hardcoded model list** (which would rot as new models ship). "Is this a thinking model?" is answered, cached per model, by: (1) **the runtime itself** — Ollama's `/api/show` reports a `thinking` capability (this is how `minimax-m3:cloud` is detected, with no name match); (2) **response-learning** — any reasoning trace, or empty-content-with-tokens, marks that model thereafter (works for any provider); (3) a small name hint as last resort. Non-reasoning models stay byte-identical, and the name-agnostic empty-content retry is the final backstop. Validated on `gpt-oss:20b-cloud`, `glm-5.2:cloud`, and `minimax-m3:cloud` (all 0% empty on the first call).
> The previously-proposed "switch to Ollama's native `/api/chat`" was a **dead end** — `/api/chat` with `think:false` *still* returns empty content for gpt-oss (it ignores it), and it would have added a fragile second code path.
## What's new in 1.12.0
Step #7 — the final step — of the [MCP modernization roadmap](./docs/audits/mcp-completeness-2026-05.md): ClarifyPrompt now speaks **A2A (Agent-to-Agent)**, so other agents can call it to compile prompts. **stdio stays the default; nothing about existing setups changes.**
Set `CLARIFYPROMPT_TRANSPORT=a2a` and ClarifyPrompt comes up as a discoverable A2A peer on Node's built-in `http` (the only new dependency is the official `@a2a-js/sdk`, which itself pulls just `uuid`):
| Endpoint | Purpose |
|---|---|
| `GET /.well-known/agent-card.json` | **Agent card** — discovery: identity, capabilities, the `compile-prompt-for-platform` skill |
| `POST /a2a` | A2A **JSON-RPC 2.0**: `message/send`, `message/stream` (SSE), `tasks/get`, `tasks/cancel`, … |
| `GET /health` | Liveness probe |
```bash
CLARIFYPROMPT_TRANSPORT=a2a CLARIFYPROMPT_HTTP_PORT=3000 npx clarifyprompt-mcp
# → card: http://127.0.0.1:3000/.well-known/agent-card.json
# → a2a: POST http://127.0.0.1:3000/a2a (message/send · message/stream)
```
The whole roadmap pays off here — one incoming A2A message flows through the same compose pipeline, and the primitives built in earlier steps map straight onto A2A semantics:
- **Compile** — a `message/send` with the raw prompt (plain text, or JSON `{ prompt, platform?, category?, … }`) returns a task whose **artifact** carries the optimized prompt (text) plus the full structured compose result (data).
- **Streaming** (1.10.0 progress → A2A) — `message/stream` emits `status-update` events as each pipeline stage runs, then the artifact, over **SSE**.
- **Cancellation** (1.10.0 AbortSignal → A2A) — `tasks/cancel` aborts the in-flight compose within milliseconds and reports a terminal `canceled` state.
- **Clarification** (1.9.0 elicitation → A2A) — clarify is **off by default** for one-shot peers; opt in with `pre_clarify: 'auto' | 'always'` and an ambiguous prompt pauses the task in A2A's first-class **`input-required`** state with the questions (readable text + structured data). Answer on the same task and it compiles.
Configure the public base URL advertised in the card with `CLARIFYPROMPT_A2A_BASE_URL` (handy behind a proxy); port/host are shared with `streamable-http`. New deterministic `npm run test:a2a` battery drives card discovery, a live compile, the clarify round-trip, and SSE streaming.
## What's new in 1.11.0
Step #6 of the [MCP modernization roadmap](./docs/audits/mcp-completeness-2026-05.md): a pluggable **transport factory** — ClarifyPrompt can now serve over **Streamable HTTP**, the runway toward A2A and remote MCP hosts. **stdio stays the default; nothing about existing setups changes.**
### Transports
Set `CLARIFYPROMPT_TRANSPORT`:
| Value | Behaviour |
|---|---|
| `stdio` (default) | One server over stdin/stdout — exactly as before |
| `streamable-http` | MCP Streamable HTTP over Node's built-in `http` (no new deps): stateful sessions (`mcp-session-id`), SSE streaming, a `/health` probe |
| `a2a` | Serve as an **A2A (Agent-to-Agent)** peer — agent card, JSON-RPC + SSE (see 1.12.0 above) |
HTTP knobs (in `streamable-http` / `a2a` mode): `CLARIFYPROMPT_HTTP_PORT` (3000), `CLARIFYPROMPT_HTTP_HOST` (`127.0.0.1` — localhost-only by default), `CLARIFYPROMPT_HTTP_PATH` (`/mcp`, streamable-http only).
```bash
CLARIFYPROMPT_TRANSPORT=streamable-http CLARIFYPROMPT_HTTP_PORT=3000 npx clarifyprompt-mcp
# → POST http://127.0.0.1:3000/mcp · GET http://127.0.0.1:3000/health
```
Tool/resource registration moved into an exported `createServer()` factory: stdio gets one server, **streamable-http gets one per session** (the SDK-recommended, GHSA-safe pattern — never shares a server across HTTP clients). New deterministic `npm run test:http` battery drives a full HTTP session.
## What's new in 1.10.0
Step #5 of the [MCP modernization roadmap](./docs/audits/mcp-completeness-2026-05.md), stable core: **`compose_prompt` is cancellable and reports live progress**. Model-agnostic, opt-in, fully back-compat.
### Cancellation
An `AbortSignal` is plumbed through the entire LLM path (`simpleGenerate` → `chat` → `fetch`, combined with the per-call timeout) and every engine stage. When a client sends `notifications/cancelled` for a `compose_prompt` call, the in-flight model request aborts immediately and the revise loop stops at the next stage boundary — instead of running every iteration to completion. The signal reaches `fetch` regardless of which model/provider is configured.
### Progress
Include a `progressToken` in the `compose_prompt` request `_meta` and the server emits `notifications/progress` at each stage (clarify / optimize / ground / critique) with a monotonic counter and a human message like `optimizing prompt [iter 2/3]`. Hosts can show a live status on a long multi-iteration compose. No token → no notifications, zero overhead.
### Why not MCP tasks (yet)
Roadmap #5 named the MCP **tasks** API. It's still `experimental/` in the SDK ("may change without notice"), its reference is ~600 lines, and no current client speaks the `tasks/*` protocol — so a full implementation would be unusable off-by-default code today. The real value (cancellable + progress-reporting compose) is delivered here on **stable** primitives; the experimental async-task wrapper is deferred to land with #7 (A2A), which the `AbortSignal` groundwork here already sets up. New deterministic `npm run test:cancel` battery locks the behavior.
## What's new in 1.9.0
Step #4 of the [MCP modernization roadmap](./docs/audits/mcp-completeness-2026-05.md): **`clarify_with_user` can elicit answers through the host's native form UI**. Opt-in, fully back-compat.
### Interactive clarification
Pass `elicit: true`. On a client that supports MCP elicitation, the clarifying questions become a real form:
- each question is a field, `options` become enum dropdowns, and each `suggestedAnswer` is the field default (one-click accept);
- the user answers inline; the engine returns `answers: [{ question, dimension, answer, usedSuggested }]` with `elicited: true`.
Without `elicit`, on a non-capable client, or if the round-trip errors, the tool returns the same raw-questions JSON it always has — every existing caller is unaffected. `decline` / `cancel` are surfaced via `elicitationAction`.
This turns clarification from "here's a JSON blob of questions, you render it" into a first-class interactive moment in hosts like Claude Desktop. The mapping lives in a small pure module (`src/engine/clarification/elicit.ts`), reusable by `compose_prompt`'s pre-clarify stage later. New deterministic `npm run test:elicit` battery (pure helpers + a live mock-client round-trip) locks it.
## What's new in 1.8.0
Step #3 of the [MCP modernization roadmap](./docs/audits/mcp-completeness-2026-05.md): the engine's read surfaces become **browseable resource templates** with **argument autocompletion**. No tool or engine behavior changes.
### Resource templates
Four templates join the static `clarifyprompt://categories`, each backed by an existing engine getter:
| URI template | What it reads |
|---|---|
| `clarifyprompt://platforms/{category}/{id}` | One platform's full config — `resources/list` enumerates all 60+ as individual URIs |
| `clarifyprompt://traces/{date}` | Optimization-trace summary index for a UTC day |
| `clarifyprompt://packs/{id}` | One loaded knowledge pack's metadata |
| `clarifyprompt://memory/facts/{scope}` | Live remembered facts under a scope |
MCP hosts with a resource browser (Claude Desktop, Cursor) now get a navigable tree instead of a single static blob.
### Autocomplete
`completion/complete` resolves the template variables: `{category}` → the 7 category ids, `{id}` → platform ids scoped by the chosen `{category}`, `{date}` → days with traces, pack ids, memory scopes. (MCP completion applies to prompt args + resource-template variables only — not tool inputs; ClarifyPrompt registers no prompts, so it lives on the templates.)
### Capabilities
The server now advertises `resources` (with templates) and `completions` at `initialize`. New deterministic `npm run test:resources` battery locks the surface.
## What's new in 1.7.1
Patch fixing [#3](https://github.com/LumabyteCo/clarifyprompt-mcp/issues/3): a **silent empty optimized prompt** from models whose answer didn't land in `content`.
- **Reads all three thinking-channel field names** (`reasoning` / `thinking` / `reasoning_content`) — fixes DeepSeek / qwen-thinking and similar.
- **Retries once, then fails loudly** when content is empty regardless of any thinking field. This covers the real issue #3 case: gpt-oss harmony output over Ollama's `/v1` shim generates tokens (`completion_tokens > 0`) but returns `content: ""` with no thinking field. The engine now degrades to the original prompt + a surfaced `error` instead of returning blank.
- **Genuinely recovering** gpt-oss harmony output (via Ollama's native `/api/chat`) was tracked as a follow-up — **resolved in 1.12.1**, which proved the `/api/chat` path a dead end and fixed the actual root cause (a `max_tokens` floor + `reasoning_effort` for reasoning models; see the 1.12.1 notes above).
- **New deterministic `npm run test:thinking`** battery locks the regression with mocked responses (no live cloud dependency).
Verified: `test:thinking`, reasoning battery (gpt-oss degrades loudly; the genuine reasoner `kimi-k2-thinking:cloud` still returns real content), integration, day2, evals, wire.
## What's new in 1.7.0
Step #2 of the [MCP modernization roadmap](./docs/audits/mcp-completeness-2026-05.md): the entire tool surface migrated off the deprecated `server.tool()` shorthand (removed in SDK 2.0) onto `server.registerTool()`. **No engine behavior changes; full back-compat.**
### What hosts get
- **Titles** — every tool has a human-readable display name ("Forget a fact", not `memory_forget`).
- **Behavior annotations** — all 23 tools declare `readOnlyHint` / `destructiveHint` / `idempotentHint` / `openWorldHint`. The three destructive tools (`memory_forget`, `unload_pack`, `unregister_platform`) are flagged for confirmation UIs; the seven read-only inspectors are flagged safe-to-call-freely; the seven tools that reach the network (LLM / embeddings / web search) carry `openWorldHint: true`.
- **Structured output** — every tool declares an `outputSchema` and returns `structuredContent` alongside the JSON text. Schemas are permissive by design (all-optional, passthrough) — they document the shape without ever rejecting engine output.
### Back-compat
Text content is byte-identical for every tool — including the three array-returning `list_*` tools, whose text stays a bare array while `structuredContent` wraps it in an object per the MCP spec. Error returns unchanged. Verified: wire 7/7, integration 9/9, day2, 26/27 evals with zero output-validation errors.
### Found during verification
[#3](https://github.com/LumabyteCo/clarifyprompt-mcp/issues/3) — cloud `gpt-oss` thinking-channel responses can yield an empty `optimizedPrompt` (remote API change exposing a pre-existing field-name gap in `client.ts`; fix targeted for 1.7.1).
## What's new in 1.6.8
Housekeeping release closing the loops the 1.6.5→1.6.7 cascade opened. **No engine code, MCP tool surface, platform, or env-var changes.**
### Changed
- **CI matrix now tests Node 24** (current active LTS, EOL Apr 2028) alongside 18/20/22 across Ubuntu + macOS. The matrix previously tested two EOL Node versions but not the current LTS at all. Verified before merge that the native deps (`better-sqlite3` + `sqlite-vec`) load and function on Node 24.16.0 in a toolchain-free `node:24-slim` container. `engines` stays `>=18` — maximum compatibility, and we test what we claim.
- **Publish runner moved Node 20 → 22**, keeping an EOL runtime off the release-critical path (matches the Dockerfile base).
### Process
- **New ship-check `CP-13 — lockfile regeneration safety`** encodes the lesson from the 1.6.5→1.6.6→1.6.7 cascade: a single `npm install --package-lock-only` silently dropped 4 of 5 `sqlite-vec` platform binaries (broke Linux CI) *and* pulled a within-caret `better-sqlite3` bump that dropped Node 20 prebuilds (broke the Docker build). The check mandates full `npm install` on dep changes, a lockfile diff for dropped platform deps + native-dep version jumps, and a local slim-Docker load gate. Dogfooded on this release.
## What's new in 1.6.7
Dockerfile patch. **No engine code, MCP tool surface, platform, or env-var changes.**
### Fixed
- **`CI / docker build` failed on 1.6.6** with `npm error gyp ERR! find Python`. Root cause: `better-sqlite3@12.10.0` ([released 2026-05](https://github.com/WiseLibs/better-sqlite3/releases/tag/v12.10.0)) explicitly removed prebuilt binaries for Node.js v20 and v23 because Node 20 reached EOL in April 2026. The 1.6.6 lockfile regen pulled 12.10.0 within the `^12.9.0` caret, and `node:20-slim` doesn't have Python + a C++ toolchain to compile from source. Bumped the Dockerfile base to `node:22-slim` — current active LTS, still has working prebuilts.
- The non-Docker CI build matrix (Node 18 / 20 / 22 across macOS + Ubuntu) still passes because regular runners can compile-from-source as fallback. Only the slim Docker image stumbles.
### Verified locally
`docker build` → green. Container can `require('better-sqlite3')` + `require('sqlite-vec')` cleanly. All 5 `sqlite-vec` platform binaries still in `package-lock.json` (1.6.6's fix held).
## What's new in 1.6.6
Lockfile + harness patch following 1.6.5. **No engine code, MCP tool surface, platform, or env-var changes.** Ships the MCP-completeness audit doc.
### Fixed
- **`package-lock.json` lost 4 of 5 `sqlite-vec` platform binaries during the 1.6.5 SDK bump.** My local `npm install --package-lock-only` retained only the maintainer's `sqlite-vec-darwin-arm64` binary. `npm ci` on CI's Ubuntu runners failed with `no such module: vec0` because `sqlite-vec-linux-x64` wasn't in the lock. End-user `npm install clarifyprompt-mcp@1.6.5` was **never affected** (the npm tarball doesn't ship a lockfile; users resolve platforms at install time). Regenerated with full `npm install` so all 5 platforms (`darwin-arm64`, `darwin-x64`, `linux-arm64`, `linux-x64`, `windows-x64`) are back.
- **Eval harness HTML report writer crashed on ERRORED entries** (`evals/run.mjs:729`). The pre-existing renderer assumed every non-skipped, non-filtered run had an `evaluation.checks` field, but errored runs carry an `error` field instead. Added an explicit errored-status branch — the harness now degrades gracefully and exits cleanly even when fixtures error.
### Bundled docs
- **[`docs/audits/mcp-completeness-2026-05.md`](./docs/audits/mcp-completeness-2026-05.md)** — diagnostic audit of the engine's MCP surface against the current SDK + spec. Tool-by-tool registration table, resource gap analysis, SDK feature delta (1.12 → 1.29 → 2.0-alpha), capability declarations, transport refactor sketch, A2A feasibility note, and a sequenced 7-step modernization roadmap. The artifact behind next-session planning. No engine changes prescribed inline.
### Numbers
- 5 sqlite-vec platforms in lockfile (was 1). `npm audit --production`: 0 vulnerabilities (unchanged). Tools: 23 (unchanged). Eval fixtures: 30 (unchanged).
## What's new in 1.6.5
Security patch. **No engine code changes, no MCP tool surface changes, no platform changes, no env-var changes.**
### Fixed
- **CVE-2026-0621** — ReDoS in `@modelcontextprotocol/sdk`'s `UriTemplate` regex (patched in SDK `1.25.2`). The previous `^1.12.1` floor *allowed* vulnerable resolutions on stale npm caches; bumped to `^1.29.0` so the floor itself is patched.
- **GHSA-345p-7cg4-v4c7** — Shared server/transport instances leak cross-client response data (patched in SDK `1.26.0`). Not exploitable in practice for ClarifyPrompt (one host = one server instance) but the vulnerable code is now out of the dependency graph entirely.
- **7 transitive vulnerabilities** (2 moderate, 5 high) in the SDK's bundled HTTP-transport substack (`hono`, `express-rate-limit`, `fast-uri`, `ip-address`, `path-to-regexp`, `qs`, `@hono/node-server`). Cleared via `npm audit fix`. Never affected runtime — ClarifyPrompt is stdio-only and doesn't load the HTTP transport — but they were noise in users' `npm audit` reports and made the install look unsafe.
### Numbers
- `npm audit --production` → **0 vulnerabilities** (was 2 SDK CVEs + 7 transitive).
- `package-lock.json`: **net −336 lines** (the old caret was pulling in heavy unused HTTP-transport ancillaries; the fix swapped them for slimmer alternates).
- Tools: 23 (unchanged). Platforms: 60+ (unchanged). Eval fixtures: 30 (unchanged).
- Wire test + integration battery + day2 + reasoning + 29/30 evals pass against the new floor on local Ollama. The one eval fail (`analyzer-creative-media`) is a pre-existing qwen-coder-7b classifier flake — verified SDK-independent by stash-reverting and re-running.
### Why the floor bump matters
`^1.12.1` was misleading documentation — caret resolution was actually pulling SDK `1.27.1` for any fresh `npm install` since early 2026. The floor bump aligns the declared baseline with what `npm` was already doing for most users while guaranteeing the floor for users on stale caches. It also positions us for the eventual `2.0.0-alpha` migration when that line stabilizes (the modern SDK deprecates `.tool()` / `.prompt()` / `.resource()` shorthand registration in favor of `registerTool()` / `registerPrompt()` / `registerResource()` with title + outputSchema + annotations).
## What's new in 1.6.4
Docs + process patch. **No engine, MCP tool, or platform changes** — but a meaningful cleanup of the pack-distribution model.
### Pack registry consolidated back into the engine repo
`LumabyteCo/clarifyprompt-packs` (the separate community-pack registry created in `1.3` with the right principle but at the wrong scale) has been **archived**. Its three starter packs already lived in this repo's [`packs/`](./packs/) folder; the registry was meant to be the canonical home but in practice everything always shipped from here via the npm tarball. The drift caught up: `higgsfield-creative-handbook` shipped in `1.6.2` and never made it to the registry, even though the registry's own README told users to fetch packs from there.
Net result of 1.6.4:
- **Single source of truth.** `packs/*.md` knowledge packs + `packs/platforms/*.yaml` platform configs all live in `clarifyprompt-mcp` and ship in the npm tarball.
- **New top-level [Knowledge packs](#knowledge-packs) section** in this README explains the loading model (`load_knowledge_pack({source: "<url-or-path>", scope: ...})`), the three starter packs + Higgsfield, the scope semantics, and how to contribute.
- **New [`packs/README.md`](./packs/README.md)** — pack authoring guide (frontmatter schema, chunk boundaries, quality bar). Lifted from the archived registry so the content isn't lost.
- **Tombstone redirect on the archived repo.** Anyone visiting `clarifyprompt-packs` lands on a banner pointing here.
### When does the split come back?
When there's a forcing function: a community PR queue on packs alone, pack count >20, or divergent licensing/governance. Until then the maintenance cost of keeping two repos in sync wasn't paying for an audience that hadn't materialized.
### Numbers
- **Tools:** 23 (unchanged).
- **Platforms:** 60+ (unchanged).
- **Bundled knowledge packs:** 4 (`anthropic-brand-voice`, `higgsfield-creative-handbook`, `nextjs-14-best-practices`, `sox-compliance`) — same as 1.6.2/1.6.3, just newly canonical.
- **Eval fixtures:** 30 (unchanged).
- **Tarball size:** unchanged from 1.6.3.
## What's new in 1.6.3
Patch. The 1.6.2 CI tag-push run surfaced two real issues — fixed here without changing any engine code.
### Fixed
- **`evals/fixtures/28-context-includes-git-state.yaml`** previously asserted `git_branch_present: true`, but GitHub Actions checks out in detached-HEAD mode where `bundle.git.branch` is correctly `undefined` (only the SHA + recent commits are populated). Relaxed to assert `bundle_has_git: true` only — that's what's actually invariant across local + CI environments.
- **`evals/fixtures/17-critique-strong-prompt-accepts.yaml`** asserted `verdict: accept` + `overall_score_min: 7` on a strong prompt. gpt-4o-mini's judge calibrates stricter than qwen2.5-coder:7b's, and occasionally returned a malformed `overall` field that the parser defaulted to 0 → verdict=reject. The fixture's *real intent* is to verify engine wiring (5+ dimensions, the standard dimension names present, no harness error) — not to compare judge calibration across models. Dropped the verdict + tight score assertions; kept the wiring-level checks.
- **README Glama badge** swapped from inline `<img>` (sometimes broken via GitHub's camo proxy) to a shields.io text-link badge that's stable across all rendering surfaces.
### Notes
- **No engine code changes.** No new MCP tools (still 23). No platform changes (still 60+). No env-var changes.
- **Eval baselines unchanged on local Ollama.** This is a CI-specific hardening — local runs against qwen-coder-7b produced the same results before and after.
- **The CI publish-gate failure that appeared on the v1.6.2 tag push** was downstream of the eval failure (`Wait for evals workflow` step blocked publish). Now that the underlying fixtures don't false-fail on gpt-4o-mini + detached-HEAD CI, the publish gate clears too.
## What's new in 1.6.2
Patch. Two additive ships, both no-code-changes from the engine's perspective:
### Higgsfield creative-handbook knowledge pack
`packs/higgsfield-creative-handbook.md` — a community-style markdown pack documenting Higgsfield's actual conventions: model-selection rules (which of the 13 models for which use case), Soul ID character-training workflow, camera-move vocabulary, prompt-structure pattern (long-form prose, not keyword tags), multi-reference editing, Marketing Studio modes, common pitfalls (don't translate Midjourney flags verbatim), output specs.
Load it explicitly:
```
load_knowledge_pack source="https://raw.githubusercontent.com/LumabyteCo/clarifyprompt-mcp/main/packs/higgsfield-creative-handbook.md"
```
…or, since it ships in the npm tarball, point at the installed copy. The Context Curator grounds Higgsfield-targeted prompts in this pack's chunks automatically via semantic retrieval. See the [Knowledge packs](#knowledge-packs) section for the full loading + scoping model.
### `npm run matrix` — multi-model eval matrix runner
`evals/matrix.mjs` runs `npm run eval` sequentially against N models and stitches the results into one side-by-side HTML (`evals/matrix.html` by default). Lights up the model-class-gated fixtures (`shape-small-local-model` / `shape-mid-tier-model` / `shape-reasoning-model`) that single-model runs skip, and exposes deltas like "qwen-7b fails analyzer-creative-media but gpt-4o-mini passes it" in a glance.
```bash
npm run matrix -- --models qwen2.5-coder:7b-instruct-q4_K_M,gpt-oss:20b-cloud,glm-5.2:cloud
```
Outputs a dark-themed table — rows = fixtures, columns = models, cells = pass / fail / skip / errored with tooltips showing which checks failed.
Companion fix: `evals/run.mjs` gains a `--json-out <path>` flag that writes structured per-model results (matrix.mjs uses it; CI agents can use it too).
### Numbers
- **No tool surface change.** Still 23 MCP tools.
- **No platform count change.** 60+ platforms (`packs/platforms/*.yaml` unchanged).
- **30 → 30 fixtures** (no new fixtures; matrix is tooling, not coverage).
- **Tarball grows ~10 KB** for the knowledge pack. `evals/matrix.mjs` is NOT in the tarball — it's a maintainer/contributor tool, not a runtime artifact.
## What's new in 1.6.1
Patch release. Adds **[Higgsfield](https://higgsfield.ai)** as a target platform in both `image` and `video` categories. No code changes — pure YAML platform-pack additions and one eval fixture.
Higgsfield is a multi-model creative platform that exposes its own MCP server at `https://mcp.higgsfield.ai/mcp`. Inside one connection you get:
- **Image**: Soul 2.0, Soul Cinema, Soul Cast (character-consistent), Flux 2, Seedream 5, Nano Banana Pro, GPT Image 2
- **Video**: Cinema Studio, Sora 2, Veo 3.1, Kling 3.0, WAN 2.6, Seedance 2.0
- **Workflows**: Soul ID character training, Lipsync Studio, UGC Factory, Marketing Studio, virality_predictor
The 1.6.1 ClarifyPrompt platform entries surface Higgsfield's model identifiers and prompt-style conventions (long-form natural-language prose; composition + lighting + textures + mood; up to 4K images / 15 s video / Soul ID for character consistency) as syntax hints to the curator.
**Recommended pattern:** install both `clarifyprompt-mcp` AND Higgsfield's MCP in your client (Claude Desktop / Cursor / AI Butler / Claude Code). Use `optimize_prompt(platform: 'higgsfield', ...)` or `compose_prompt(platform: 'higgsfield', ...)` to compile, then pass the compiled prompt to Higgsfield's `generate_image` / `generate_video` tool. MCPs compose at the client; ClarifyPrompt stays at the "compile" layer.
29 → 30 eval fixtures. Same MCP tool surface as 1.6.0 (23 tools, 1 resource). No env-var changes.
## What's new in 1.6.0
Four targeted additions across the engine's four pillars (memory / agentic / models / context), each shipped behind real eval fixtures. **3 new MCP tools** (23 total). Fully back-compat with 1.5.x — no removed tools, no removed fields, no required env-var changes.
### Memory — explicit fact CRUD (`memory_remember`, `memory_forget`, `memory_list_facts`)
Before 1.6, facts only entered persistent memory via *reflection on `save_outcome`* — implicit, LLM-extracted, after-the-fact. 1.6 adds the explicit path:
- **`memory_remember`** — directly insert a `(subject, predicate, object)` triple with explicit confidence. Source tagged `user:explicit`. Auto-embedded for future semantic retrieval.
- **`memory_forget`** — soft-delete (bi-temporal `invalidated_at`) a fact by id. Idempotent: re-forgetting an already-invalidated fact is a no-op and returns `success: false` cleanly.
- **`memory_list_facts`** — list live facts in a scope (default `user`), optionally filtered by predicate. Sorted by most-recently-observed.
This closes the obvious UX gap where the engine could only learn from *outcomes* — now users can say "remember I prefer X" directly.
### Agentic — `compose_prompt`'s new `max_iterations` revise loop
`compose_prompt` used to revise *once* (the critique's `improvedPrompt` replaced the optimization, if the verdict wasn't `accept`). 1.6 adds a loop:
```json
{ "prompt": "...", "post_critique": true, "auto_revise": true, "max_iterations": 3 }
```
Each iteration after the first re-runs `optimize` + `critique` on the previous iteration's improved prompt. Stops at `verdict=accept`, no improvedPrompt to feed back, or the cap. `pre_clarify` only runs once (no point re-asking on a rewrite). The response includes a new `iterations` field showing how many fired. Hard cap of 5 to prevent cost runaways.
### Models — per-stage model routing
Each compose stage can now target a different model:
```json
{
"prompt": "...",
"clarify_model": "qwen2.5-coder:7b-instruct-q4_K_M",
"optimize_model": "claude-sonnet-5",
"critique_model": "gpt-4o-mini"
}
```
Run clarify on a cheap local model, optimize on the big-budget frontier model, critique on the cheap judge. The override flows through every layer — `optimization.metadata.model` and `critique.judgeModel` in the response reflect the actual model that ran each stage.
### Context — git-state + environment signals
Two new signal collectors feed the Context Curator:
- **`bundle.git`** — current branch, short SHA, dirty flag, last 5 commit titles. Lets the engine ground prompts in "what you're iterating on" without you spelling it out. Detected via `git rev-parse` / `git status` / `git log`; fails soft when cwd isn't a repo.
- **`bundle.environment`** — `nowIso` / `weekday` / `timezone` (IANA from `Intl.DateTimeFormat`). Helps with time-sensitive prompts ("send this email tomorrow"). Pure JS, never fails.
Both are low-utility candidates in the curator (won't dominate budget) but surface as grounding sources when relevant.
### Eval coverage
23 → **29 fixtures** (6 new):
- 24 `memory-remember-persists` / 25 `memory-forget-invalidates` — Me1 CRUD round-trip
- 26 `compose-loop-iterates` — A1 loop infrastructure (new `iterations_min` / `iterations_max` checks)
- 27 `compose-per-stage-models-honored` — M1 per-stage routing (new `optimization_model_eq` / `critique_model_eq` checks)
- 28 `context-includes-git-state` / 29 `context-includes-environment-time` — C1 + C4 signals (new `bundle_has_git` / `bundle_has_environment` / `git_branch_present` checks)
Local baseline on `qwen2.5-coder:7b`: **25 passed / 1 failed / 3 skipped / 97% avg**. The lone failure remains the persistent `analyzer-creative-media` model-class signal (untouched).
## What's new in 1.5.2
The first release where CI's eval gate (against `gpt-4o-mini`) drove the diff. Three real fixes that the gate caught the moment we wired in the `OPENAI_API_KEY` secret:
- **Memory store now supports any embedding dimension** ([#2](https://github.com/LumabyteCo/clarifyprompt-mcp/issues/2)). The persistent vec table was hardcoded to 768 dims (the nomic-embed-text default), so anyone configuring `EMBED_MODEL=text-embedding-3-small` (1536), `voyage-3` (1024), `embed-english-v3.0` (1024), or any non-768 model would hit `Dimension mismatch: expected 768, got N` on the first `memory_search` call. The store now derives the table name from the embedder's actual dimension and creates the dim-specific table at boot. Existing 768-dim installs are unaffected.
- **`LLM_TIMEOUT_MS` env-var override** on the LLM client. Default stays at 30s; users on slow hosted models can bump it. The eval workflow uses 120s for `gpt-4o-mini`.
- **Eval harness hardened** — no longer crashes when a tool throws an exception (the SDK returns plain-text error responses; the harness used to `JSON.parse` them and die). One bad fixture no longer tanks the whole run.
- **Live evals badge.** The `evals.yml` workflow runs on every push to main. The `[![evals]](...)` badge at the top of this README is its real-time status. Currently green at 20/0/3 · 100% on `gpt-4o-mini`.
No new MCP tools. No env-var surface changes (only an added optional `LLM_TIMEOUT_MS`). Fully back-compat with 1.5.x.
## What's new in 1.5.1
A patch release on top of 1.5.0. Pure docs + ship-process improvements; **runtime behavior is identical to 1.5.0**.
- **README marketing surfaces refreshed** — the 1.5.0 release shipped with the README still on 1.4.0 in three places (headline blockquote, "What's new in X" heading, "cumulative through X" annotation). Every other version surface (`package.json`, `package-lock.json`, `server.json`, `src/index.ts`, `CHANGELOG`) was correct, but the prose drifted because nothing automated touched it. 1.5.1 fixes that.
- **Two new ship-check audits** — `CP-11` (README marketing-surface coherence) hard-fails if any of the three above don't reference the current `package.json#version`. `CP-12` (Platform-pack format validity) parses every `packs/platforms/*.yaml` and asserts schema validity. CP-11 was promoted to the user-scoped (cross-project) ship-check skill the same day, so future projects benefit too.
- **No code changes.** No new MCP tools. No new env vars. Same tarball anatomy as 1.5.0 plus a few hundred bytes of CHANGELOG.
## What's new in 1.5.0
**Built-in platforms become declarative.** The 58+ hardcoded TypeScript platform arrays move to `packs/platforms/*.yaml` — adding a built-in platform is now a YAML edit, not a TS edit. The TypeScript layer becomes a runtime loader with a hardcoded fallback table. Malformed YAML can never soft-brick the server.
```
packs/platforms/
chat.yaml 9 platforms
code.yaml 9
document.yaml 8
image.yaml 10
music.yaml 4
video.yaml 11
voice.yaml 7
README.md contributor docs
```
To add a new built-in platform: append an entry to the relevant category file, run `npm run build`, open a PR. No TS edit required. Custom-platform-via-runtime (`register_platform`) still works identically for user-installed platforms.
- **Memory-layer eval coverage.** The eval harness now supports `setup: [{tool, args}, ...]` — a list of MCP tool calls executed BEFORE the main `input`. Two new fixtures use it: one loads a knowledge pack inline and verifies the chunk surfaces in `grounding.sources` after the embed → store → retrieve → curate → ground pipeline; the other proves vector-search ranking quality. **23 fixtures total** (was 20 in 1.4.0).
- **Test infrastructure modernization.** The integration + Day-2 test batteries used to assert literal version strings (`1.3.0`, `16 tools`) and broke on every bump. Now they read `EXPECTED_VERSION` from `package.json` and assert presence of a tool *set* rather than a tool *count*. Future bumps don't break the tests.
- **Adoption materials.** `docs/adoption/` ships with copy/paste-ready Show HN body, Reddit posts, Twitter thread, awesome-mcp-servers PR template, and catalog submission specs (mcp.so, Smithery, mcp-get, PulseMCP, modelcontextprotocol/servers).
- **One new runtime dep:** `js-yaml` promoted from devDependency for the platform loader (~200 KB).
- **Same MCP tool surface as 1.4.** 20 tools, 1 resource. No new tools; no removed tools; result shapes unchanged.
## Previously in 1.4.0 — the composable pipeline
Four core operations as first-class MCP tools that compose. Use any tool standalone, or run the whole chain in one call:
```
┌─────────────┐ ┌─────────────────────┐ ┌──────────────┐
│ clarify │ → │ ground OR optimize │ → │ critique │
│ (optional) │ │ (core) │ │ (optional) │
└─────────────┘ └─────────────────────┘ └──────────────┘
one call = compose_prompt(prompt, [sources], post_critique, auto_revise, ...)
```
- **`clarify_with_user`** — Given an ambiguous draft, returns 1–3 targeted clarifying questions, each with a `suggested_answer` you can accept verbatim, optional 2–4 quick-pick `options`, and a `dimension` tag (audience/scope/format/length/tone/constraints/goal/platform). Short-circuits with `clarificationNeeded: false` on confident, well-formed prompts so it pipelines cleanly in front of `optimize_prompt` without a per-call latency tax.
- **`ground_prompt`** — The strict, retrieval-augmented variant of `optimize_prompt`. Caller-provided sources are pinned at the **highest** priority — above project rules, above pinned instructions — and tracked individually in the trace as `user-source:N`. Strict mode: zero non-empty sources → error, no silent fall-through. Per-source body cap (4000 chars) so a single huge paste can't dominate the budget.
- **`critique_prompt`** — LLM-as-judge. Scores a candidate prompt 0–10 across 5 default dimensions (clarity, specificity, intent_alignment, format_fitness, length_appropriateness) — or your own criteria — with per-dimension rationale + concrete suggestions, an overall score, and a verdict (`accept` / `revise` / `reject`). Below `revise_threshold` (default 7.0) it also returns an `improvedPrompt` you can drop in. Use it pre-flight ("is this prompt good enough for the expensive model?"), postmortem ("was the prompt the cause?"), or to A/B-pick the best of N optimization variants.
- **`compose_prompt`** — One MCP call runs the canonical pipeline. Auto-decides the ground vs. optimize branch from whether you passed `sources`. `pre_clarify: 'auto' | 'always' | 'never'`. `post_critique: true` adds a judge pass. `auto_revise: true` replaces `final_prompt` with the rewrite when the verdict isn't `accept`. Returns a per-stage `stages` audit array so the caller sees exactly what ran.
- **Eval harness v0** — Deterministic regression tests under `evals/`. 20 YAML fixtures cover analyzer, shape, intent-overlay, grounding, clarify, critique, ground, and compose surfaces. `npm run eval` produces a console summary + self-contained dark-themed HTML report. Multi-model matrix is just bash: run `LLM_MODEL=... npm run eval -- --report-path evals/report-X.html` per model.
- **CI-gated evals (opt-in)** — When `OPENAI_API_KEY` is set as a repo secret, the eval harness runs in CI against `gpt-4o-mini` as a release gate. Off by default; nothing leaves your machine without the secret.
- **5 new MCP tools** (20 total). `optimize_prompt` also gains a `userProvidedSources` injection point — both `ground_prompt` and `compose_prompt` use it under the hood, but it's available directly if you want explicit control without the strict-mode validation.
> Carried over from 1.3: persistent memory + knowledge packs + reflective learning. The curator continues to score and fit grounding sources into the target model's remaining window. `explain_last_curation` still gives you a per-call breakdown of selected vs. rejected candidates with reasons.
## What's in the box (cumulative through 1.15.0)
- **Context Engine** — auto-gathers workspace rules (`CLAUDE.md`, `AGENTS.md`, `.cursorrules`, `.clinerules`, `clarify.md`), detects frameworks and languages from `package.json` and sibling manifests, tracks an active file excerpt, and maintains a per-session ring buffer of recent optimizations **and their outcomes**.
- **Unified `PromptAnalyzer`** — one LLM call produces `{ category, intent, recommendedMode, confidence }` together. 10 intents: `production-code`, `brand-voice`, `stakeholder-comm`, `data-extract`, `creative-media`, `technical-spec`, `analysis`, `quick-draft`, `exploration`, `unknown`. Intent beats surface keywords on ambiguity.
- **Target-model-aware prompt shaping** — system prompt, `maxTokens`, and `temperature` adapt to the downstream LLM's context window and the resolved intent. Small local models get a compact prompt; Claude/GPT-4/Gemini get the full richness.
- **Grounding Context (single, priority-ordered)** — user pinned instructions → project rules → active file → prior accepted examples → web search → workspace metadata → target-model hints → custom platform instructions → built-in syntax hints. No more parallel context silos.
- **Session retrieval (save_outcome)** — the caller reports `accepted | edited | rejected` per optimization; similar accepted outputs in the same session get injected as few-shot examples into future similar prompts. Backed by persistent memory (SQLite + sqlite-vec), so accepted outcomes survive restarts.
- **Local JSONL tracing** — every optimization writes a structured trace line (now with `shape`, `groundingSources`, `error` fields) to `$CLARIFYPROMPT_HOME/traces/YYYY-MM-DD.jsonl`. **Nothing is uploaded.** Toggle via `CLARIFYPROMPT_TRACE=off`.
- **Unified `$CLARIFYPROMPT_HOME`** — one env var for everything ClarifyPrompt writes. Legacy `CLARIFYPROMPT_CONFIG_DIR` / `CLARIFYPROMPT_DATA_DIR` still work (deprecation hint, silenceable).
- **Three transports** — `stdio` (default), `streamable-http` (MCP over Node `http`, stateful sessions + `/health`), and `a2a` (an **Agent-to-Agent** peer: agent card, JSON-RPC `message/send` + SSE `message/stream`, task cancellation, `input-required` clarification). One `CLARIFYPROMPT_TRANSPORT` env var; stdio behavior is byte-identical to before.
- **60+ platforms, 7 categories, custom platforms** — the original core is unchanged and fully backward-compatible.
- **Any LLM, any provider.** One code path works with **any OpenAI-compatible API** — Ollama (local + cloud), LM Studio, vLLM, OpenAI, Google Gemini, xAI Grok, Groq, Mistral, DeepSeek, Cohere, Perplexity, Together, Fireworks, OpenRouter — plus **Anthropic Claude** directly. Reasoning models (`o1/o3/o4`, `deepseek-reasoner`, `gpt-oss`, `*-thinking`) are auto-detected and given a larger token budget so they actually produce content. [See 15+ pre-configured provider examples below](#provider-examples).
- **Apache-2.0, forever.** Open-source core, no relicensing.
## Quick Start
### With Docker
Pull the published image from GitHub Container Registry (multi-arch: `amd64` + `arm64`, with signed provenance + SBOM):
```bash
docker pull ghcr.io/lumabyteco/clarifyprompt-mcp:latest
```
All config is passed at **run time** — nothing is baked into the image, so the image is safe to share and contains no secrets:
```bash
# stdio (for MCP hosts that launch the container)
docker run --rm -i \
-e LLM_API_URL=http://host.docker.internal:11434/v1 \
-e LLM_MODEL=qwen2.5:7b \
-e CLARIFYPROMPT_HOME=/data \
-v clarifyprompt-data:/data \
ghcr.io/lumabyteco/clarifyprompt-mcp:latest
# or serve over HTTP / A2A
docker run --rm -p 3000:3000 \
-e CLARIFYPROMPT_TRANSPORT=a2a -e CLARIFYPROMPT_HTTP_HOST=0.0.0.0 \
-e LLM_API_URL=http://host.docker.internal:11434/v1 -e LLM_MODEL=qwen2.5:7b \
ghcr.io/lumabyteco/clarifyprompt-mcp:latest
```
> Mount a volume at `CLARIFYPROMPT_HOME` to persist memory, traces, and packs across runs. Pass `LLM_API_KEY` / `EMBED_API_KEY` as `-e` env vars (or `--env-file`) at run time — never bake them into an image.
### With Claude Desktop
Add to your `claude_desktop_config.json`:
```json
{
"mcpServers": {
"clarifyprompt": {
"command": "npx",
"args": ["-y", "clarifyprompt-mcp"],
"env": {
"LLM_API_URL": "http://localhost:11434/v1",
"LLM_MODEL": "qwen2.5:7b"
}
}
}
}
```
### With Claude Code
```bash
claude mcp add clarifyprompt -- npx -y clarifyprompt-mcp
```
Set the environment variables in your shell before launching:
```bash
export LLM_API_URL=http://localhost:11434/v1
export LLM_MODEL=qwen2.5:7b
```
### With Cursor
Add to your `.cursor/mcp.json`:
```json
{
"mcpServers": {
"clarifyprompt": {
"command": "npx",
"args": ["-y", "clarifyprompt-mcp"],
"env": {
"LLM_API_URL": "http://localhost:11434/v1",
"LLM_MODEL": "qwen2.5:7b"
}
}
}
}
```
### With AI Butler
[AI Butler](https://github.com/LumabyteCo/aibutler) is a self-hosted
personal AI agent runtime — single Go binary, multi-channel chat, MCP
ecosystem hub. Drop ClarifyPrompt into its `mcp.servers` config and
the agent picks up all 23 tools as native capabilities, callable from
any channel (web chat, terminal, Telegram, Slack, etc.). AI Butler
discovers tools dynamically via MCP's `tools/list`, so adding /
removing tools in ClarifyPrompt updates the agent's surface
automatically — no config edits needed on the butler side.
Edit `~/.aibutler/config.yaml`:
```yaml
configurations:
mcp:
servers:
- name: clarifyprompt
command: clarifyprompt-mcp
env:
LLM_API_URL: "http://localhost:11434/v1"
LLM_MODEL: "qwen3-vl:8b"
```
Restart AI Butler. The boot log confirms the tools are wired in:

The agent enumerates the full surface on request — every tool prefixed
with `clarifyprompt.`:

> 📸 **Screenshots above are from a 1.2-era integration (11 tools).**
> Current `1.6.x` exposes **23 tools** — `optimize_prompt`,
> `clarify_with_user`, `ground_prompt`, `critique_prompt`,
> `compose_prompt`, plus the management / inspection / memory
> tools (`memory_search`, `memory_remember`, `memory_forget`,
> `memory_list_facts`, knowledge-pack tools, traces, custom
> platforms, etc.). AI Butler picks them up automatically via the
> MCP `tools/list` discovery; no config changes needed.
#### Drive the Context Engine end-to-end
You can preview what the engine *would* gather (without running the
optimization) using `inspect_context`:
![Context Engine preview — analyzer output (Category=code, Intent=production-code, Recommended Mode=detailed, Confidence=Medium), session history, and the priority-ordered grounding stack the engine would merge into the system prompt. Closing takeaway about a language mismatch the engine detected between workspace (JS) and prompt (TypeScript).](docs/screenshots/aibutler/03-inspect-context.p