io.github.jarmstrong158/context-keeper
Durable project-memory MCP: decisions, constraints, and pipelines across Claude sessions.
Open source Open in the app JSON README (API)
About
Durable project-memory MCP: decisions, constraints, and pipelines across Claude sessions.
Details
- Kind
- MCP servers
- Topic
- AI, RAG & memory
- Publisher
- jarmstrong158
- Origin
- official
- Category
- ferramentas
- Transport
- local
- Version
- 0.19.0
- Stars
- 1
- Last push
- 2026-08-17T12:58:07Z
- Repository state
- ativo
- Language
- Python
- License
- MIT
- Added
- 2026-08-29 04:00:14
- Updated
- 2026-08-29 04:00:14
- Origin id
io.github.jarmstrong158/context-keeper
README
<!-- mcp-name: io.github.jarmstrong158/context-keeper -->
# Context Keeper
_Part of the [xylem](https://github.com/jarmstrong158/xylem) stack._
Project memory for Claude. Records design decisions, pipeline flows, and constraints so Claude maintains context across conversations.
## The Problem
As conversations get long, Claude loses the "why" behind earlier decisions. New conversations start blank. This causes Claude to make changes that break established patterns — like rewriting a pipeline step it doesn't remember exists.
## The Solution
Context Keeper gives Claude 14 tools to record and retrieve structured project context:
| Tool | Purpose |
|------|---------|
| `record_entry` | Unified write tool — record a decision, pipeline, or constraint via `kind`, with per-kind fields validated server-side. Consolidates the former `record_decision`/`record_pipeline`/`record_constraint` (still dispatchable by those names for back-compat) |
| `get_context` | Retrieve relevant entries by query, tags, scope, or ID — relevance-ranked, pulls `related_to` links by default |
| `query_entries` | **Exact structured-field filtering** (status, origin, tags, scope, hardness, supersession, dates) — deterministic, no ranking; distinct from `get_context`'s relevance search |
| `get_project_summary` | Compact overview for conversation start |
| `update_entry` | Update any entry by ID |
| `deprecate_entry` | Retire an entry with reason (optional `merge_into` folds a duplicate into a survivor) |
| `prune_stale` | Find entries not verified recently |
| `get_compaction_report` | Check if last compaction lost any context |
| `verify_quality` | Scan entries for thin rationale, missing tags, isolated arcs (auto-called by PreCompact hook) |
| `export_markdown` | Regenerate `DECISIONS.md` from the decisions store — a derived, read-only projection |
| `reload_constraints` | Re-surface the constraints-only block on demand mid-session (rules refresh, not the full store) |
| `export_snapshot` | Write the whole store to a committable `.context-keeper/memory.json.gz` for sharing project memory via git |
| `import_snapshot` | Import that committed snapshot into the working store — non-destructive, auto-runs on first use when the store is empty |
| `mirror` | Sync with the optional remote store: `op="pull"` merges remote→local (newest wins), `op="backfill"` pushes local→remote. No-op if the remote is unconfigured |
All data stored as human-editable JSON files in `.context/` inside your project directory. Zero dependencies by default, semantic retrieval optional.
### How this relates to Claude Code's built-in memory
Claude Code ships [two memory mechanisms of its own](https://code.claude.com/docs/en/claude-md): **CLAUDE.md** files you write by hand, and **auto memory**, where Claude saves freeform notes to `~/.claude/projects/<project>/memory/`. Context Keeper is not a replacement for either — it sits on different ground, and the differences are the reason to run it:
| | Auto memory | Context Keeper |
|---|---|---|
| **Shape** | Freeform markdown, an index plus topic files | Typed entries (decision / pipeline / constraint) with server-validated fields |
| **Depth** | Whatever Claude writes | Schema-enforced: `problem` ≥40 chars, `why_chosen` ≥60, thin entries rejected at capture |
| **Lifecycle** | Edit or delete the file | `supersedes` (demoted, still recallable), `deprecate`, `merge_into`, drift + staleness scans |
| **Conflicts** | None | Restatement vs contradiction classified at capture, with origin-based trust precedence |
| **Retrieval** | Index loaded whole each session | Relevance-ranked within a token budget, with an abstention signal on no-answer queries |
| **Reach** | Machine-local; explicitly not shared across machines | Cross-device via the mirror, team-shared via a committed snapshot |
They compose rather than compete: auto memory is good at picking up incidental preferences with zero effort, and Context Keeper is for the decisions and rules you want structured, queryable, enforceable, and portable. Running both is fine — and with `rules_export` enabled, Context Keeper writes into the harness's own `.claude/rules/` surface rather than around it.
**Known gap, stated honestly:** subagents do not inherit the main conversation's session-start injection, so a subagent starts without the project summary. If your workflow leans on subagents, have them call `get_project_summary` explicitly — the retrieval-is-unskippable property holds for the main loop only.
## Capabilities at a glance
Context Keeper is a small, offline-first memory layer; several of its capabilities are easy to miss because they live inside existing tools rather than as separate features. The map below names them in memory-system terms:
| Capability | How Context Keeper does it |
|------------|----------------------------|
| **Procedural memory** | `record_entry(kind="pipeline")` stores ordered, dependency-aware workflows (build/deploy/data flows) with `purpose` + `when_to_invoke` — reusable "how we do X", not just facts. |
| **Deduplication** | Every `record_*` runs a word-set Jaccard pass against the store and returns `similar_entries` when a new entry restates an existing one, so duplicates are caught at capture; `deprecate_entry(merge_into=...)` then folds the duplicate's unique content into the survivor and retires it in one non-destructive step. |
| **Contradiction detection** | Those same overlaps are classified `likely_restatement` vs `likely_contradiction` (negation/antonym polarity), and a reversal raises a `contradiction_note` telling the agent to resolve the conflict rather than leave two live rules disagreeing. |
| **Quality refinement** | `verify_quality` scans for thin rationale, missing tags, legacy-schema entries, and isolated (unlinked) arcs; the PreCompact hook runs it automatically so entries get enriched before context is compressed. |
| **Supersede / decay / forget** | `supersedes` demotes-but-keeps prior decisions (recallable history); `prune_stale` surfaces unverified entries for review; `deprecate_entry` removes an entry from retrieval entirely. |
| **Origin + trust / source attribution** | Every entry records `origin` (`user` / `agent` / `import`); retrieval gives user-stated entries a trust boost and it decides the default winner when entries conflict. |
| **Anticipated queries** | `retrieval_hints` stores alternate phrasings a future session might search for, so vocabulary-mismatch queries hit without embeddings. |
| **Hybrid retrieval** | Lexical (tag + word overlap) by default; an opt-in embedding-cosine blend (`semantic.enabled`) adds vector recall, with lexical fallback when the embedder is offline. |
| **Fact-metadata query** | `query_entries` filters entries by exact predicates over structured fields (status, origin, tags-any/all, scope, hardness, supersession, dates), AND-combined and deterministic — a precise lookup path distinct from `get_context`'s fuzzy relevance ranking. |
| **Cache-friendly injection** | The session-start memory block is deterministically ordered with a stable prefix and the only per-session-volatile line (quality-scan IDs) emitted last, so an unchanged store injects byte-identical text across sessions. |
| **Path-triggered rules** | Scoped constraints project into Claude Code's own `.claude/rules/*.md` format with `paths:` frontmatter (`rules_export`), so the harness loads a rule when the agent *reads* a covered file — before an edit, with no hook involved. The `scope_guard` PreToolUse hook covers the write path for clients without rules support. |
| **Narrative + clustering** | `get_project_summary` clusters decisions by topic above a threshold and renders a compact narrative; the `DECISIONS.md` projection mirrors the store as human-readable prose. |
| **Data export / offline / privacy** | Plain JSON in `.context/` you can read, edit, grep, and commit; runs fully offline with zero required dependencies and no data leaving the machine. |
### Evaluation & benchmarks (open methodology)
The retrieval and honesty properties are measured, not asserted — the harness is in [`evals/`](evals/) and reproducible with no network required:
- **Token reduction** — session-start injection vs. dumping the full store: **97.3% / 94.1% / 85.5% / 73.3%** across four real stores ([`evals/token_reduction.py`](evals/token_reduction.py)). The meaningful property is that injected cost stays roughly flat as the store grows.
- **Retrieval quality** — 59 cases over a frozen 7-store corpus, every question written from the *problem* an entry solves rather than paraphrasing its summary. The opt-in semantic blend lifts **recall@5 from 0.42 → 0.64 and MRR 0.37 → 0.55** ([`evals/run_retrieval_eval.py`](evals/run_retrieval_eval.py)). These are *lower* than the figures published before 2026-08-05 (hit@5 80% → 93%) and that is the point: the old set was partly paraphrase-derived, so queries shared surface tokens with their targets and lexical recall came out flattered. The gain is concentrated in large, prose-heavy stores; small stores are already at 1.000 lexically.
- **Abstention** — measures whether `get_context` says "nothing relevant" instead of confabulating on no-answer queries; the 0.20 relevance floor is the highest with zero false-abstention on the eval set ([`evals/abstention.py`](evals/abstention.py)).
Every dataset, metric, and caveat is checked into the repo — see [`evals/README.md`](evals/README.md). The corpus is frozen under `evals/fixtures/corpus`, so the numbers reproduce on any clone and a regression test can pin them; `--live` runs against the real stores when you want a current read instead of a comparable one.
## v0.19: Supersession as a Signal, and One Rule for Scope
- **The store could say what was true, and not what changed.** It held 36 entries, 36 active, 0 deprecated, 0 supersedes links — while `dec-013` stated it replaced the old additive-only sync and `con-006` carried "(was additive-only)" inline. `deprecate_entry` had supported `superseded_by` since v0.10 and it had never been used once: nothing asked at the moment the answer was known, and nothing spent the link afterwards.
`record_*` now names active same-kind entries covering the same subject — `possible supersession of <id>: <summary>` — scored on shared tags and whole-component scope overlap, deliberately **not** the text overlap `similar_entries` already computes (a replacement often shares almost no wording with what it replaces while addressing exactly the same thing). It never links, never blocks the write, and never touches the older entry: a heuristic edge silently demotes a rule that may still be in force.
`get_context` prepends one compact line for the **immediate** predecessor — what it said and why it changed — because a link is only worth making if retrieval spends it. One level deep; under budget pressure the *line* is dropped and a `predecessor_id` trail is left, never the entry. The remote Worker emits the identical line, so the two transports cannot diverge.
`scripts/survey_supersessions.py` proposes backfill links **read-only** — it opens no store for writing and calls no lifecycle tool.
- **`con-011` was false.** It requires every surface deciding "does this scope cover this file" to agree; there were five surfaces with four implementations, two of which provably disagreed, and `score_entry` still used the raw substring test the rule exists to forbid — so querying `scope="hooks/"` gave a `webhooks/`-scoped entry the same boost as the real one. `scope_rules.py` is now the only implementation, importing nothing so the PreToolUse hook can load it under `con-010`.
- **Retrieval has a number, and it holds still.** A 59-case golden set with negatives and history cases, plus a regression test pinning the lexical arm. The first pin broke within a day without a line of ranking code changing — it read live stores, so recording three entries moved it. The corpus is frozen now.
- **Internals**: `server.py` 4,270 → 3,787 lines (`ranking.py`, `mojibake.py`, `quality_checks.py`, all re-exported); `verify_quality`'s seven inlined checks became a registry; the test suite split from one 5,213-line file into six themed ones.
## v0.17: Closing the Delivery Gaps
v0.16 got a rule in front of the model *before* the edit. Measuring real
stores afterwards showed two holes that compound, and one that had been open
since the beginning.
- **Truncated entries are no longer undiscoverable.** The SessionStart hook
prints the summary text and nothing else, so an entry dropped for budget
wasn't merely unsummarised — the agent had no way to learn it existed in
order to go ask for it. On the largest real store that was **116 lines of
memory silently gone**. The summary now ends with the ids it dropped, paid
for **inside** the same budget (bare ids cost a fraction of the entries they
stand for) and emitted in stable store order so the prompt-cache prefix stays
byte-identical. `summary_dropped_ids` carries the same list structurally.
The constraints floor still wins over the budget when it has to — one real
store's 19 constraints are 2429 tokens against a 2000 budget — because
losing the rules is worse than overspending. It is reported, never silent.
- **`global` scope is now a flagged quality issue.** 47% of constraints across
real projects were `global`, which excludes them from the drift check, the
`.claude/rules/` projection *and* `scope_guard` — leaving the truncating
summary as their only route to the model. Those two problems multiply.
`verify_quality` now flags `global_scope`, but **only when it can name a
concrete path** from the entry's `enforced_by` or a tag matching a real
directory. A bare "consider adding a scope" on every global rule is noise
the reader learns to skip, and some rules genuinely are global.
- **Subagents get the project rules.** `SessionStart` does not fire for
subagents and they do not inherit the parent's injected context, so every
subagent has been starting with no project memory — on a fan-out, a dozen
contributors who never read the rules. The new `subagent_start.py` hook
injects the constraints-only block via `SubagentStart`. Constraints only,
and capped: this fires once per spawned agent, so a fan-out multiplies it.
```json
"SubagentStart": [
{ "hooks": [{ "type": "command",
"command": "python /path/to/context-keeper/hooks/subagent_start.py",
"timeout": 5 }] }
]
```
- **A release can no longer half-fail in silence.** v0.16.0 published to PyPI
and the MCP registry, then the bundle workflow failed on manifest
validation — and a release missing an asset is indistinguishable from one
that never had it. A new `verify-release` workflow asserts all three
channels actually received the version, and the manifest is now validated on
every PR rather than at release time.
- **`constraint_reinject` came off the heavy import path.** Its matcher is
`""`, so it runs after *every* tool call — the hottest hook in the set — and
it was importing `server` to format a list of strings. Constraint rendering
moved to `store_paths`, so it and `subagent_start` both run within ~5ms of
Python's own startup floor. One implementation still, shared by every
surface.
## v0.16: Path-Triggered Rules, and the Rule Before the Edit
Both halves of this release come from reading [Anthropic's own memory
documentation](https://code.claude.com/docs/en/claude-md) and asking which of
its mechanics context-keeper was leaving on the table.
- **Scoped constraints project into `.claude/rules/*.md` (opt-in).** Claude
Code loads a rule file carrying `paths:` frontmatter when it **reads** a
matching file — earlier than any hook can fire, and with no hook wired at
all. `rules_export.enabled` mirrors every scoped constraint into
`.claude/rules/context-keeper/`, one file per scope, on the write path of
every constraint mutation. Same contract as the `DECISIONS.md` projection:
JSON stays canonical, the markdown is derived, regenerated whole, and never
parsed back in.
The projection owns a **subdirectory** rather than `.claude/rules/` itself,
and reaps only files carrying its own generated marker — so a deprecated
constraint's rule file disappears on the next write, while a hand-written
rule dropped in the same folder is never touched. The marker is a
block-level HTML comment, which Claude Code strips before the file enters
context: it identifies the file to the tool and to a human reader at zero
token cost.
A scope is **refused** rather than emitted when it can't become a pattern
that actually fires: glob metacharacters (`[ ] { } * ?`), because Claude Code
reads `[` as a bracket expression and one that never closes matches nothing;
and a double quote or control character, because each pattern is emitted as a
double-quoted YAML scalar and an embedded quote closes it early, killing the
whole frontmatter block — the harness then loads *no* rule from that file,
not even the patterns that were fine. Refused scopes come back in
`skipped_scopes`, because a rule file that silently never fires is worse than
no rule: it looks like coverage.
`rules_export.path` is **contained to the project**. This directory is not
merely written to, it is reaped — every marker-carrying `.md` file in it is a
deletion candidate on each regeneration — so a `../../..` or absolute path in
config would aim that deletion at a directory the user never associated with
context-keeper. Containment is also just correct: Claude Code only discovers
`.claude/rules/` inside the working tree, so an outside directory could never
load anyway. A path that escapes is refused with an error, and the constraint
write itself still succeeds — projections are derived, the JSON store is
canonical, and a misconfigured projection must never fail a record.
- **`scope_guard` now runs under PreToolUse.** It fired on PostToolUse, which
means the rule arrived *after* the write had already landed — a review note,
not a guardrail. PreToolUse `additionalContext` is injected next to the tool
result, so the constraint reaches the model while it can still act on it.
One script serves either wiring: it reads `hook_event_name` and answers with
the matching `hookEventName`, defaulting to PostToolUse when the field is
absent, so **existing installs keep working unchanged** and upgrading is a
config edit rather than a rewrite.
Opt-in `scope_guard.confirm_absolute` escalates a PreToolUse hit on an
**absolute** constraint to `permissionDecision: "ask"`, pausing for the user
instead of only annotating. Default off. The honest limit, stated plainly:
this escalates on **scope, not on violation** — nothing here reads the diff,
so it cannot know the edit actually breaks the rule.
- **The edit path got a latency budget it can't quietly lose.** Moving to
PreToolUse means the hook runs *before* the tool, so its cost lands on the
critical path of every Edit and Write rather than trailing them. Measured on
Windows, the hook cost **~142ms**, of which **~73ms was `import server`** —
`mirror` pulling in `urllib.request` → `http.client` → `email.parser`, plus
`secrets`, `usage`, `code_drift`, none of which this hook touches. Worse, a
path matching *no* constraint paid the same 142ms as a hit, and that is most
edits.
Store location and raw reads now live in `store_paths.py`, which imports
`json` and `os` and nothing else. `server.py` imports its resolution from
there rather than defining its own, so the precedence order (env var → Xylem
pointer → cwd → parent walk) has exactly one implementation — a second copy
would drift, and the copy that drifts is the one nothing executes.
**~142ms → ~69ms**, against a ~62ms bare-interpreter floor. The remainder is
Python process startup, which a shell hook cannot avoid. `TestEditPathHookCost`
now fails if an edit-path hook imports `server` again, or if `store_paths`
grows an import beyond `json`/`os` — the property is pinned, not just fixed.
Wire the hook with `"timeout": 5` (shown below) to bound the pathological
case; nothing else caps a hook that hangs.
- **Over-budget summaries are now floored and reported.** Truncating the
session-start summary popped lines from the end, and with a small enough
budget it walked straight through the constraints block and emptied the
summary — the v0.9 failure (a healthy-looking store injecting nothing) in a
different guise. Trimming now stops at the constraints block, and the
response carries `summary_truncated` / `summary_lines_dropped` /
`summary_truncation_note` when the cap bit. Those keys appear **only** when
truncation happened, so the untruncated common case stays byte-stable for
the prompt cache. Anthropic's auto-memory index errors rather than silently
dropping content past its read limit; same principle.
- **The store can now report — and repair — its own encoding damage.**
`con-008` fixed the *cause* of mojibake in v0.11: stdin defaulted to cp1252
on Windows, so a client's raw UTF-8 bytes were mis-decoded before
`json.loads` ever ran and an em-dash landed in the store as `â€"`. Forcing
UTF-8 on the transport meant no new entry was corrupted. It did nothing for
entries already written, and nothing ever looked.
That damage is invisible in the worst possible way: the text stays legible
enough that nobody re-reads the entry, so a corrupted rationale quietly
degrades every retrieval that surfaces it. `verify_quality` now flags it as
`mojibake`, and a `repair_mojibake` handler fixes it:
```bash
context-keeper repair_mojibake '{}' # dry run, writes nothing
context-keeper repair_mojibake '{"apply": true}' # repair
```
The repair is the exact inverse of the corruption and is **verified as
one** — re-applying the corruption to the candidate must reproduce the
input byte for byte, and anything failing that check is left alone. A
partial repair of someone's recorded reasoning is worse than legible
damage, because it looks fixed. `verified_at` is deliberately not
refreshed: an encoding fix is not a claim that anyone re-confirmed the
entry is still true, and resetting the staleness clock would erase the
signal `prune_stale` and the drift check exist to raise. `updated_at` *is*
bumped, so the corrected copy wins the mirror's newest-wins merge rather
than being overwritten by a corrupt remote.
Enable the projection in `.context/config.json` and backfill once:
```json
{ "rules_export": { "enabled": true } }
```
```bash
context-keeper export_rules '{}'
```
`export_rules` is deliberately **not** in `tools/list` — con-004 caps the
schema payload every client pays for at session start, and this is a one-time
backfill. Render-on-write keeps the directory current afterwards with no tool
call.
**Commit it or ignore it, but match your store.** The generated directory is
derived from `.context/`. If your working store is gitignored (the default),
gitignore `.claude/rules/context-keeper/` too — otherwise a fresh clone's first
constraint write regenerates from a store it doesn't have and reaps every
committed file. If you share memory through `export_snapshot`, committing the
rules directory is fine and gives a new teammate the scoped rules immediately.
## v0.15: Two-Way Mirror (local <-> remote)
Optional, fail-soft mirroring so a second device — e.g. a phone recording decisions on the go — can both **receive** the desktop's memory and **contribute** its own. The local `.context/` JSON store stays **canonical**; the [context-keeper-remote](https://github.com/jarmstrong158/context-keeper-remote) Cloudflare Worker is a sync surface, never the source of truth. (Distinct from the git-committed `export_snapshot` mechanism — the mirror is live and cross-device, the snapshot is a versioned bundle in the repo.)
Conflict resolution is **last-writer-wins by `updated_at` timestamp**, applied identically in both directions, so an edit made on either device converges everywhere. Neither side ever deletes — a deprecation is a status change that propagates like any other edit.
**The clock-skew caveat, stated honestly.** "Last writer" is decided by comparing ISO-8601 `updated_at` strings, and those timestamps come from *whichever machine did the write* — the desktop's wall clock for a local edit, the Worker's for a remote one. Two edits to the *same* entry made within the clock skew between those machines (realistically sub-second, cross-device) can therefore misorder: the copy stamped later isn't guaranteed to be the one written later, so the guard can pick the wrong winner. There is no logical clock or vector clock to break the tie — just wall time. The blast radius is bounded (a single pull compares only the remote's own timestamps, against one clock, so intra-remote ordering is exact; skew only bites when a local write races a remote one on the same id), and it is never silent: whenever newest-wins overwrites a copy that differed in substance, the losing version is appended to `.context/.mirror_conflicts.json` (see below) so you can reconcile by hand.
- **Mirror out (local -> remote).** After every write (`record_entry`/`record_*`, `update_entry`, `deprecate_entry`) the entry is pushed via the remote's `upsert_entries` MCP tool (one call per kind). `upsert_entries` preserves the incoming id and **replaces** an existing remote copy only when the pushed entry's `updated_at` is newer — so an edit or deprecation actually overwrites the stale remote copy instead of being skipped. If the remote is unreachable the entry is queued to `.context/.mirror_queue.json` (deduped to its latest state) and flushed on the next successful push. A push failure **never** blocks or fails the local write.
- **Mirror in (remote -> local).** `pull_remote` calls the remote's `query_entries` and merges **every** returned row **by timestamp** — a remote entry whose id exists locally overwrites the local copy only when the remote's `updated_at` is newer; a newer local copy is kept and pushes back on its next write. The `.context/.mirror_watermark` is only a bookkeeping hint now: every fetched row goes through the per-entry newest-wins merge, so a remote entry is **never** dropped merely because its timestamp falls at or below the watermark (an earlier build filtered on the watermark *before* merging, which could silently lose a phone-recorded entry under cross-device clock skew — the exact scenario the mirror exists for). Wired into the SessionStart hook (so desktop sessions start with phone-recorded entries present) and exposed through the `mirror` MCP tool as `op="pull"`.
- **Conflicts are preserved, not lost.** When either direction overwrites a copy that differed in substance (not just timestamps), the losing version is appended to `.context/.mirror_conflicts.json`. Last-writer-wins has already resolved which copy is live; this is the audit trail of what it replaced. No resolution UI — the record is there if you need to reconcile by hand.
- **Backfill** (`mirror` with `op="backfill"`) pushes the entire local store to the remote (one `upsert_entries` call per kind) — for seeding a fresh remote. Idempotent: an equal-or-older re-push is skipped server-side.
- **Collision-safe IDs (Option B: random suffix).** Two stores minting sequential ids independently would collide — the desktop and the Worker both hand out `dec-013` for *different* decisions, and an upsert keyed by `(project, id)` would then let one silently overwrite the other. Fix: when mirroring is enabled, new ids get a short random hex suffix — `dec-013-a7f3`. The number still leads (sortable, greppable); **old ids are never rewritten**; and with mirroring *off* ids stay bare `dec-013` (single writer, no collision possible).
- **Zero new dependencies** (stdlib `urllib` only), **no secrets in code**: the remote URL contains the auth token as its final path segment (`/mcp/<token>`) and comes from an env var only.
Enable by setting one env var where the MCP server runs:
```bash
CONTEXT_KEEPER_REMOTE_URL=https://context-keeper-remote.<acct>.workers.dev/mcp/<AUTH_TOKEN>
# CONTEXT_KEEPER_REMOTE_TIMEOUT=5 # optional per-request seconds
```
With no `CONTEXT_KEEPER_REMOTE_URL` set, every mirror path is a silent no-op — behavior is identical to pre-v0.15.
**Transport:** stateless JSON-RPC over Streamable HTTP. Each write is one `POST` to the `/mcp/<token>` URL (`tools/call` → `upsert_entries`); the pull op (`mirror` with `op="pull"`) calls the remote's `query_entries`. No initialize/session handshake (the server is stateless); the response is a single `application/json` body.
## v0.14: Dedup Merge (`deprecate_entry(merge_into=...)`)
Capture-time detection already caught near-duplicates (`similar_entries` with a `likely_restatement` relation), but *resolving* one was a manual two-step: deprecate the duplicate, then `update_entry` the original to fold in anything it was missing. v0.14 collapses that into one atomic, non-destructive operation.
- **Opt-in param on the existing tool, not a new tool.** `deprecate_entry(id=<dupe>, reason=..., merge_into=<survivor>)` folds the duplicate's unique content into the survivor, then deprecates the duplicate with `superseded_by=<survivor>`. When `merge_into` is absent, `deprecate_entry` behaves exactly as before — byte-for-byte.
- **Additive and non-destructive.** The survivor can only *gain* content: list fields (`tags`, `retrieval_hints`, `related_to`, `constraints`/`constraints_created`) are unioned, and empty text fields are backfilled from the duplicate — a non-empty field on the survivor is **never** overwritten. The duplicate isn't hard-deleted; it stays on disk as a deprecated entry pointing at the survivor, so the merge is fully auditable and reversible.
- **Same-type, single-write, validated first.** Merge requires both entries to be the same type (so their schemas line up), resolves and validates the target before any write, and mutates both entries in one file write so the two updates can't clobber each other. A bad `merge_into` (missing target, cross-type, or self) errors cleanly and deprecates nothing.
- **Roots held.** Explicit and agent-invoked (like every other lifecycle tool), zero new dependencies, no LLM call, deterministic. It streamlines the restatement workflow the capture loop already prescribes rather than adding a background process.
```jsonc
// dec-002 restates dec-001 — merge and retire it in one call
{ "id": "dec-002", "reason": "Restatement of dec-001", "merge_into": "dec-001" }
// -> dec-001 gains dec-002's unique tags/hints/related_to + any text it lacked;
// dec-002 becomes deprecated with superseded_by = dec-001
```
## v0.13: Structured Field Query (`query_entries`)
`get_context` answers *"what's relevant to what I'm working on?"* — it ranks by relevance, blends optional semantics, and flags low-relevance results with an abstention signal. That's the right tool for fuzzy recall, but the wrong one when you already know the exact field values you want. `query_entries` fills that gap: **deterministic filtering over the structured fields that already exist on every entry**, no ranking and no abstention.
- **Exact predicates, AND-combined:** `types`, `status` (active/superseded/deprecated), `origin` (user/agent/import), `tags_any`, `tags_all`, `scope` (exact, case-sensitive), `hardness` (absolute/advisory), `supersedes` / `superseded_by`, and the same `since` / `before` temporal filters as `get_context`. Every predicate is a hard match over an existing field — a query either matches or it doesn't.
- **No relevance, no confabulation.** Results come back in stable natural-ID order with no score and no `min_relevance` floor — an empty result set is a real, honest answer, not an abstention message. The abstention machinery is for fuzzy text queries; a structured predicate doesn't need it.
- **Same store, same budget.** It reuses the exact store-reading and entry-serialization paths `get_context` uses, and packs the matched set into the same token budget (default 4000, `token_budget` per call), so a broad query can't dump the store — `matched_entries` vs `entries_returned` and a `budget_truncated` flag tell you if the cap clipped anything.
- **Additive and self-contained.** Zero new dependencies, stdlib only, no embeddings and no LLM call — pure in-memory filtering over JSON already on disk. `get_context`, the semantic blend, the scoring, and every hook are untouched; default behavior of every existing tool is byte-for-byte unchanged.
One deliberate difference from `get_context`: **`query_entries` applies no default status filter**, so `superseded` and `deprecated` entries *are* returned unless you pass `status`. `get_context` always hides deprecated entries; the structured tool lets you ask for them on purpose.
**Examples:**
```jsonc
// Absolute constraints scoped to the hooks/ directory
{ "types": ["constraints"], "hardness": "absolute", "scope": "hooks/" }
// User-stated decision that superseded dec-005
{ "origin": "user", "supersedes": "dec-005" }
// Active pipelines tagged "release"
{ "types": ["pipelines"], "status": "active", "tags_any": ["release"] }
// Everything a user asserted this month, across all types
{ "origin": "user", "since": "2026-07-01" }
```
## v0.12: Contradiction Detection + Cache-Stable Injection
- **Restatement vs contradiction, at capture time.** The similar-entry pass already caught heavy overlaps; now it classifies each one. Two dependency-free signals — negation asymmetry ("X is required" vs "X is *not* required") and antonym polarity ("always" here / "never" there, "enable" / "disable") — label a match `likely_restatement` or `likely_contradiction`. A restatement nudges you to merge; a contradiction raises a `contradiction_note` telling the agent to resolve which rule is current (`deprecate_entry` with `superseded_by`) instead of silently leaving two live rules that disagree. Advisory only, and only evaluated on pairs Jaccard already flagged as overlapping — the write always proceeds. Zero new dependencies, no LLM call, no added tokens at record time.
- **Cache-stable session-start injection.** The injected memory block is ordered so its large stable portion — constraints, decisions, pipelines, and the fixed capture guidance — forms a prefix that repeats byte-for-byte across sessions when the store hasn't changed, while the one volatile line (the quality scan's flagged IDs) is emitted last. This keeps the memory block inside the model's cacheable prompt prefix rather than busting the cache each session. It also *reduces* tokens rather than adding them.
## v0.11: Mid-Session Constraint Re-Injection (opt-in)
The SessionStart hook injects your constraints once, at turn one. As a long
session fills with tool output, those rules scroll out of the model's working
attention and effectively decay — the model can violate a constraint it was
briefed on an hour ago simply because it is buried. v0.11 re-surfaces the
constraints **during** a long session, not just at the start.
Two ways in, both **constraints-only** — they re-inject the exact
Absolute/Advisory block SessionStart shows, and nothing else from the store
(no decisions, no pipelines). It's a lightweight rules refresh, not a second
full dump.
- **`reload_constraints` tool** — returns the current constraints block on
demand. Always available; call it whenever you want the rules back in
context.
- **`constraint_reinject.py` hook (PostToolUse)** — **opt-in, default off.**
When enabled, it counts tool calls per session and re-injects the
constraints block every *N* calls (`every_n_tools`, default 25) via
`additionalContext`.
**What triggers it, honestly.** The automatic path is the **PostToolUse**
hook — that surface *is* injected into the model, and its firing rate tracks
tool-output volume, which is the thing actually burying the rules. It is **not
a timer**: an MCP server has no wall-clock inside the context window, so
re-injection is driven by counting tool calls, not elapsed seconds. It is
**not PreCompact** either — PreCompact stdout is shown only to the user, never
injected into the model (the compaction boundary is already re-covered by the
SessionStart hook, which re-fires with source `compact`).
**Default behavior is unchanged.** With no config (or `enabled: false`), the
hook is inert and SessionStart works exactly as before. Enable it in
`.context/config.json`:
```json
{ "constraint_reinjection": { "enabled": true, "every_n_tools": 25 } }
```
## v0.10: Abstention + Supersession-as-Ranking
Two ideas adapted from studying [Curion](https://github.com/geanatz/curion), kept dependency-free:
- **`get_context` can now say "I don't have anything relevant."** Previously it always returned its top-scored entries — but the composite score banks ~55 points from recency/status/origin regardless of relevance, so a query with *no* relevant memory silently got a confident-looking result. Measured confabulation was **100%** on no-answer queries (`evals/abstention.py`). Now the response carries `top_relevance` and, when the top entry's tag/text relevance falls below `min_relevance` (config, default 0.20), `no_confident_match: true` with guidance telling the agent not to present the entries as established fact. It **annotates, never suppresses** — weak matches are still returned, so the vocabulary-mismatch recall that `retrieval_hints` and the semantic blend preserve survives. 0.20 is the highest floor with zero false-abstention on the eval set.
- **Supersession as a ranking signal, not just a filter.** `record_decision` accepts `supersedes: [ids]`: the prior decisions become `superseded` — **demoted in ranking but still recallable** ("why did we change from X?"), distinct from `deprecate_entry` which removes an entry from retrieval entirely. Superseded entries are skipped by `prune_stale`/`verify_quality` (they're intentional history, not stale work) and marked `**SUPERSEDED** by dec-NNN` in the `DECISIONS.md` projection.
Deliberately *not* adopted from Curion: its LLM-controller architecture (an API call on every store and recall). context-keeper stays zero-dependency and offline by default.
## v0.9: Topic Clustering, More Embedding Backends + a Bug the Measurement Caught
- **Critical fix: empty session-start injection for large stores.** The summary truncation loop evaluated the *original* text in its condition, so any store whose summary exceeded the token budget (~30+ entries) silently popped every line and injected an **empty** summary at session start. Found while measuring token reduction: a 78-entry store was injecting ~0 tokens of memory. Now truncates correctly to budget.
- **Topic clustering.** Above 8 decisions, `get_project_summary` groups decisions by their most-frequent shared tag instead of one flat list — a 59-decision store reads as a dozen topics.
- **OpenAI-compatible embeddings.** `semantic.api: "openai"` points the semantic blend at any `/v1/embeddings` endpoint — LM Studio, llama.cpp server, or OpenAI itself (`api_key_env` names the env var holding the key). Ollama stays the default; same fail-safe lexical fallback. nomic task prefixes now apply only to nomic models.
- **Trust-aware conflict guidance.** `similar_entries` matches now carry each entry's `origin`, and the guidance states the precedence: user-stated overrides agent-inferred overrides imported.
- **Token-reduction measurement** (`evals/token_reduction.py`), run against four real stores:
| store | active entries | full store (tokens) | injected at session start | reduction |
|---|---|---|---|---|
| balatron | 78 | ~75,277 | ~2,057 | 97.3% |
| clark | 55 | ~35,445 | ~2,102 | 94.1% |
| context-keeper | 13 | ~5,692 | ~828 | 85.5% |
| conductor | 9 | ~1,538 | ~411 | 73.3% |
Baseline = dumping every active entry into context; injected = the `get_project_summary` output the SessionStart hook prints. Honest caveat: the summary is budget-capped (default 2000 tokens), so for large stores part of the reduction is by construction — the meaningful property is that injected cost stays flat as stores grow.
- **Six more MCP clients documented** (OpenCode, Copilot CLI, Antigravity, OpenClaw, Hermes, pi/oh-my-pi) — see Other MCP clients below.
## v0.8: DECISIONS.md Projection (render-on-write)
Opt-in: mirror the decisions store into a human-readable `DECISIONS.md` at the project root. Enable in `.context/config.json`:
```json
{ "markdown_export": { "enabled": true, "path": "DECISIONS.md" } }
```
- **Render-on-write.** Every tool call that mutates a decision (`record_decision`, `update_entry`, `deprecate_entry`) regenerates the entire file from `decisions.json` after the JSON write and before the tool returns — so a subsequent `git commit` captures both in the same commit. Deliberately *not* a git/PostToolUse hook: rendering after the commit snapshot would reintroduce drift.
- **JSON stays canonical; markdown is derived and read-only.** The file is regenerated whole every time — never appended to, merged, or parsed back in. Hand edits are not preserved; a regenerated projection has no drift surface.
- **`export_markdown` tool** regenerates on demand (optionally to a custom `path`), so existing repos can backfill without enabling the flag.
- Pure stdlib string formatting; default behavior with the flag off is byte-for-byte unchanged.
Born from field use: Balatron's `DECISIONS.md` was kept in sync with the store by hand, one mirror-edit per commit. This automates that convention.
## v0.7: Anticipated Queries, Origin Trust, Timeline Filters
- **`retrieval_hints`** (all `record_*` tools): 2-4 alternate phrasings a future session might search for — synonyms, symptom descriptions, error messages. Indexed for both lexical and semantic retrieval, so vocabulary-mismatch queries ("value network diverging" vs. "value head saturating") can hit without embeddings. The zero-dependency complement to the semantic blend.
- **`origin` + trust weighting** (all `record_*` tools): entries record who authored them — `user` (explicitly stated), `agent` (inferred from the session), or `import` (backfilled). Retrieval scoring gives user-stated entries a trust boost over agent-inferred, which outrank imports. Pre-v0.7 entries score as `agent`, preserving their relative order.
- **`since` / `before` on `get_context`**: temporal filters against each entry's verified/created timestamp — "what did we decide this month" is now a query.
## v0.6: Capture-Time Guardrails
- **Scoped constraint injection.** New `scope_guard.py` hook (PostToolUse on `Edit|Write|NotebookEdit`): the moment the agent edits a file covered by a constraint's `scope`, that constraint is injected into context via `additionalContext`. Session-start injection briefs the model once at turn one; this enforces the rule at the exact moment it's about to matter. Each constraint fires at most once per session.
- **Similar-entry surfacing at record time.** `record_*` now compares the new entry against the store (word-set Jaccard, threshold configurable via `similar_threshold`) and returns `similar_entries` when existing entries overlap heavily — catching restatements and contradictions at capture instead of relying on MMR to mitigate duplicates at retrieval. Advisory only: the write always proceeds.
## v0.5: Data Integrity + Retrieval Fixes
- **Atomic writes.** Entry files are written to a temp file and swapped in with `os.replace`, so a crash mid-write can no longer leave a truncated JSON file behind.
- **Corrupt-store protection.** If an entry file exists but can't be parsed, `record_*`/`update_entry`/`deprecate_entry` now refuse to write (previously a corrupt file read as empty, and the next record silently replaced your entire history with one entry). Read-only tools still degrade gracefully.
- **`update_entry` enforces the schema.** Structured fields (`why_chosen`, `problem`, `reason`, `purpose`, ...) are min-length validated on update too, so entries can't be hollowed out after recording.
- **Better budget packing.** `get_context` skips entries that don't fit the token budget and keeps packing smaller ones, instead of stopping at the first oversized entry.
- **Fresh compaction reports.** The SessionStart hook now runs the snapshot comparison itself (SessionStart fires with source `compact` immediately after compaction — before any Stop), so the injected report is never one compaction stale. It also injects a one-line quality-scan nudge, which is the model-visible surface for `verify_quality` (PreCompact stdout is only shown to the user, not the model).
- **Semantic layer shipped in the package** (`semantic_index.py` was missing from the wheel/sdist), with batched embedding requests and one fewer HTTP round-trip per query.
## v0.4: Structured Rationale + Arc Linking
Earlier versions used a single freeform `rationale` field. In practice, agents wrote one-line summaries instead of full reasoning — defeating the point. v0.4 fixes this three ways:
1. **Schema-enforced depth.** `record_decision` requires `problem` (min 40 chars), `why_chosen` (min 60 chars), and accepts optional `what_we_tried` and `tradeoffs`. `record_pipeline` requires `purpose`. `record_constraint` enforces `reason` ≥ 40 chars and accepts optional `triggering_incident`. Thin entries are rejected server-side with field-specific guidance — the lazy path no longer produces a useful entry.
2. **Arc linking via `related_to`.** Every entry can reference IDs of related entries. `get_context` traverses these links by default (depth=1), so when you retrieve one decision the rest of its arc comes along. Connective tissue survives across sessions.
3. **Quality verification.** A new `verify_quality` tool scans for legacy entries, thin reasoning, missing tags, and isolated entries (tag overlap with no `related_to`). The `PreCompact` hook calls it automatically and surfaces flagged entries so they can be enriched before context is compressed.
Legacy entries (pre-v0.4) stay valid — they're never auto-rejected, just flagged by `verify_quality` for optional enrichment. The deprecated `rationale` parameter still works on `record_decision` for backward compatibility (it auto-maps to `why_chosen`), but `problem` is still required.
## Install
Two ways to install, depending on your client. Claude Desktop users get the
one-click bundle; everything else uses the standard stdio server.
### Option A — Claude Desktop one-click bundle (.mcpb)
Context Keeper ships as an [MCPB desktop extension](https://github.com/anthropics/mcpb):
a single `.mcpb` file you install without touching any config.
1. Download `context-keeper-<version>.mcpb` from the
[Releases page](https://github.com/jarmstrong158/context-keeper/releases).
2. Double-click it (or drag it into Claude Desktop → Settings → Extensions).
3. When prompted, choose a **Storage directory** — the folder where your project
memory lives (a `.context/` subfolder of readable JSON is created there). Then
enable the extension.
That's it — no `pip`, no JSON editing. The bundle is stdlib-only Python, so it has
no third-party dependencies to install. (Claude Desktop provides the Python
runtime for `.mcpb` python extensions; you need Python available for it to launch.)
The bundle is built reproducibly from this repo with `scripts/build-mcpb.sh`, and
CI attaches it to each version's GitHub Release automatically.
### Option B — pip + stdio (Claude Code, Cursor, Codex, any MCP client)
```bash
pip install context-keeper-mcp
```
#### Claude Code
```bash
claude mcp add --scope user context-keeper -- python /path/to/context-keeper/server.py
```
#### Claude Desktop (manual config)
Prefer editing config by hand instead of the `.mcpb` bundle? Add to your `claude_desktop_config.json`:
```json
{
"mcpServers": {
"context-keeper": {
"command": "python",
"args": ["/path/to/context-keeper/server.py"],
"env": {
"CONTEXT_KEEPER_PROJECT": "/path/to/your/project"
}
}
}
}
```
#### Other MCP clients (Cursor, Codex CLI, Gemini CLI, Windsurf, ...)
The server is a standard stdio MCP server, so any MCP-capable client can use it — the hooks are Claude Code extras, not requirements. Point your client's MCP config at `python /path/to/context-keeper/server.py` and set `CONTEXT_KEEPER_PROJECT`:
**Cursor** (`~/.cursor/mcp.json` or per-project `.cursor/mcp.json`) and **Windsurf** (`~/.codeium/windsurf/mcp_config.json`) use the same shape as Claude Desktop:
```json
{
"mcpServers": {
"context-keeper": {
"command": "python",
"args": ["/path/to/context-keeper/server.py"],
"env": { "CONTEXT_KEEPER_PROJECT": "/path/to/your/project" }
}
}
}
```
**OpenAI Codex CLI** (`~/.codex/config.toml`):
```toml
[mcp_servers.context-keeper]
command = "python"
args = ["/path/to/context-keeper/server.py"]
env = { "CONTEXT_KEEPER_PROJECT" = "/path/to/your/project" }
```
**Gemini CLI** (`~/.gemini/settings.json`) uses the same `mcpServers` JSON shape as Cursor above.
**GitHub Copilot CLI** (`~/.copilot/mcp-config.json`) and **oh-my-pi** (`mcpServers` config) use the `mcpServers` shape with `"type": "stdio"`:
```json
{
"mcpServers": {
"context-keeper": {
"type": "stdio",
"command": "python",
"args": ["/path/to/context-keeper/server.py"],
"env": { "CONTEXT_KEEPER_PROJECT": "/path/to/your/project" }
}
}
}
```
**OpenCode** (`opencode.json`):
```json
{
"mcp": {
"context-keeper": {
"type": "local",
"command": ["python", "/path/to/context-keeper/server.py"],
"environment": { "CONTEXT_KEEPER_PROJECT": "/path/to/your/project" }
}
}
}
```
**Antigravity** (`~/.gemini/config/mcp_config.json` or workspace `.agents/mcp_config.json`) and **OpenClaw** (`openclaw.json`) use the `mcpServers` shape with `command`/`args`, same as Copilot above.
**Hermes** (`~/.hermes/config.yaml`):
```yaml
mcp_servers:
context-keeper:
command: "python"
args: ["/path/to/context-keeper/server.py"]
env:
CONTEXT_KEEPER_PROJECT: "/path/to/your/project"
```
Without the Claude Code hooks you lose automatic session-start injection and edit-time constraint guards — call `get_project_summary` at conversation start and `record_*` as you work instead (the tool descriptions prompt for this).
Set `CONTEXT_KEEPER_PROJECT` to the root of your project. If omitted, the server resolves the project directory in this order:
1. **`CONTEXT_KEEPER_PROJECT`** env var (explicit opt-in — trusted)
2. **cwd** if it already contains a `.context/` directory
3. **Walk parent dirs** from cwd looking for an existing `.context/` (git-style discovery — finds your project when the server is launched from any subdirectory of it)
4. Otherwise: refuse, and `record_*` returns an "unresolved project" error
Steps 2 and 3 only resolve to directories that **already** contain `.context/`. The server never creates one implicitly, so you can never accidentally pollute a parent directory by launching from the wrong place. Pass `project_dir` explicitly to any tool to force-create a new project.
## How It Works
### Recording Context
When you make a design decision:
```
You: Let's use JSON files instead of SQLite for storage.
Claude: [calls record_entry(kind="decision") with summary, problem, why_chosen,
alternatives, and optionally what_we_tried + tradeoffs + related_to links]
```
When you establish a workflow:
```
You: The deploy pipeline is: run tests, build, push to registry, deploy.
Claude: [calls record_entry(kind="pipeline") with ordered steps]
```
When you set a rule:
```
You: Never run Conductor from source. Always use the exe.
Claude: [calls record_entry(kind="constraint") with rule, reason, and hardness=absolute]
```
### Retrieving Context
At conversation start, the SessionStart hook injects the project summary (and any compaction-discrepancy report) directly into context — no tool call required, so retrieval can't be skipped on a task-focused first turn. `get_project_summary` remains callable on demand. Before making changes, Claude calls `get_context` with relevant tags to check for conflicts.
### Relevance Scoring
Without embeddings or external services, Context Keeper scores entries using:
- **Tag match** — overlap between query and entry tags
- **Text match** — query words found in summary/rationale/rule text
- **Recency** — recently verified entries score higher
- **Status** — active entries prioritized over superseded
Results are capped by a configurable token budget (default: 4000 tokens).
## Claude Code Hook Setup
Context Keeper includes hooks that inject project memory at session start, remind Claude to capture after every git commit, snapshot your context before Claude Code compaction, and detect if anything was lost afterward.
Add to your Claude Code hooks config (`~/.claude/settings.json`):
```json
{
"hooks": {
"PreCompact": [
{
"matcher": "",
"hooks": [
{
"type": "command",
"command": "python /path/to/context-keeper/hooks/pre_compact.py"
}
]
}
],
"Stop": [
{
"matcher": "",
"hooks": [
{
"type": "command",
"command": "python /path/to/context-keeper/hooks/post_compact.py"
}
]
}
],
"SessionStart": [
{
"matcher": "",
"hooks": [
{
"type": "command",
"command": "python /path/to/context-keeper/hooks/session_start.py"
}
]
}
],
"PreToolUse": [
{
"matcher": "Edit|Write|NotebookEdit",
"hooks": [
{
"type": "command",
"command": "python /path/to/context-keeper/hooks/scope_guard.py",
"timeout": 5
}
]
}
],
"PostToolUse": [
{
"matcher": "Bash",
"hooks": [
{
"type": "command",
"command": "python /path/to/context-keeper/hooks/commit_capture_reminder.py"
}
]
},
{
"matcher": "",
"hooks": [
{
"type": "command",
"command": "python /path/to/context-keeper/hooks/constraint_reinject.py"
}
]
}
]
}
}
```
The `constraint_reinject.py` entry is only active when
`constraint_reinjection.enabled` is set in `.context/config.json` (default
off) — wiring it up is harmless until you opt in. Its matcher is `""` (every
tool call) so the per-session counter advances on all activity.
Replace `/path/to/context-keeper` with the actual install path. Set `CONTEXT_KEEPER_PROJECT` env var if your project isn't in the current working directory.
**Windows users:** Use forward slashes (`C:/Users/.../context-keeper/hooks/pre_compact.py`) or double-escaped backslashes in JSON. Single backslashes get mangled by the shell.
The hooks form a complete capture-and-retrieval loop:
- **SessionStart** — imports the server's own handlers and prints the project summary (plus any compaction-discrepancy report and a one-line quality-scan nudge) straight to stdout, which Claude Code injects into context at turn one. It also runs the post-compaction snapshot comparison itself before reading the report — SessionStart fires with source `compact` immediately after compaction, before any Stop hook, so this keeps the injected report fresh. This replaces the older approach of printing an instruction to *call* the tools — a request that reliably lost to a task-focused first turn since the tools are deferred. Stays silent when the project has no `.context/` yet, and emits ASCII-only output so it cannot crash on Windows cp1252 stdout
- **PostToolUse (Bash)** — fires after every Bash tool call; when the command contains `git commit`, it injects a reminder to record the matching decision/constraint/gotcha **in the same work cycle**. A commit is the single best capture trigger — it's the exact moment something became real enough to persist in version control. Born from field use: during incident-heavy sessions the agent batched capture "for later," and the user had to ask "update context keeper" three times in one night while a dozen commits shipped
- **PreToolUse (Edit|Write)** — `scope_guard.py`: when the agent is about to write a file covered by a constraint's `scope` (e.g. a constraint scoped to `hooks/` and an edit to `hooks/session_start.py`), that constraint is injected via `additionalContext` **before the write executes**. Session start briefs the rules; this puts them in front of the model at the moment of edit. Once per constraint per session. It was wired under PostToolUse through v0.15 — that still works (the hook detects its own event and defaults to PostToolUse), but the rule arrived after the edit had landed, which made it a review note rather than a guardrail. Optional `scope_guard.confirm_absolute` escalates an absolute-constraint hit to a user confirmation prompt; note it triggers on scope, not on an actual violation
- **PostToolUse (any tool)** — `constraint_reinject.py`: **opt-in, default off.** When `constraint_reinjection.enabled` is set, it counts tool calls per session and re-injects the constraints-only block every `every_n_tools` calls via `additionalContext`, so rules injected at session start don't decay as tool output buries them. PostToolUse is chosen deliberately: it's a model-visible surface (unlike PreCompact) and its firing rate tracks tool-output volume. Not a timer — an MCP server has no wall-clock in the context window
- **PreCompact** — snapshots all active `.context/` entries and runs a quality scan (`verify_quality`), printing flagged entries (thin reasoning, missing tags, isolated arcs) to the transcript. Note: PreCompact stdout is user-visible only — Claude Code does not inject it into the model's context, which is why the model-visible quality nudge lives in the SessionStart hook instead
- **Stop** — safety-net run of the same snapshot comparison SessionStart performs, in case the session ends without a new session starting (idempotent — skips if the snapshot hasn't changed since last comparison)
This closes the capture loop: SessionStart injects retrieval at turn one, the commit reminder anchors capture to the moment changes land, PreCompact is the pre-compression safety net, and Stop handles integrity checking. Retrieval is unavoidable; capture is now *prompted at the right moment* rather than left to the agent's discretion mid-task.
## Data Storage
```
your-project/
.context/
decisions.json # Design decisions with rationale
pipelines.json # Multi-step workflows
constraints.json # Rules and invariants
config.json # Token budget, stale threshold
embeddings.json # Semantic-retrieval vector cache, keyed by entry text (auto-generated, only when semantic enabled)
compaction_snapshot.json # Pre-compaction snapshot (auto-generated)
compaction_report.json # Post-compaction diff report (auto-generated)
reinject_state.json # Per-session tool counter for constraint re-injection (auto-generated)
scope_guard_state.json # Per-session record of already-injected scoped constraints (auto-generated)
.mirror_queue.json # Queued mirror-out writes pending a reachable remote (auto-generated)
.mirror_watermark # Newest remote timestamp already pulled (auto-generated)
.mirror_conflicts.json # Substance-differing versions overwritten by newest-wins (auto-generated)
hook.log # Hook activity log
mirror.log # Mirror (local<->remote) activity log (auto-generated)
.claude/
rules/
context-keeper/ # Scoped-constraint rules projection, one file per