io.github.jaimenbell/rag-mcp
Minimal RAG-over-a-corpus MCP retrieval: search_knowledge returns cited chunks. Local embeddings.
Open source Open in the app JSON README (API)
About
Minimal RAG-over-a-corpus MCP retrieval: search_knowledge returns cited chunks. Local embeddings.
Details
- Kind
- MCP servers
- Topic
- AI, RAG & memory
- Publisher
- jaimenbell
- Origin
- official
- Category
- ferramentas
- Transport
- local
- Version
- 0.1.0
- Last push
- 2026-09-03T17:47:00Z
- Repository state
- ativo
- Language
- Python
- License
- MIT
- Added
- 2026-08-29 04:00:14
- Updated
- 2026-08-29 04:00:14
- Origin id
io.github.jaimenbell/rag-mcp
README
# rag-mcp
[](https://github.com/jaimenbell/rag-mcp/actions/workflows/ci.yml)
> A minimal, honest **RAG-over-a-corpus MCP retrieval tool**. One tool,
> `search_knowledge(query, k, doc_class=None)`, that embeds a query, vector-searches a local
> corpus, and returns passages **with citations** (source + heading + chunk index) so answers
> are traceable.
Built to slot into the [mcp-factory](https://github.com/jaimenbell/mcp-factory) manifest model.
Fully local + **$0** (no paid embedding API).
## Why it's safe to put in front of a real corpus
- **Cited** - every hit carries `source` + `heading` + `chunk_index`.
- **Auth-scoped** - results are confined to the configured corpus root; sources that escape it
(absolute paths, `..` traversal) are refused.
- **Fail-soft** - a down or empty store returns a *structured error*, never an exception that
crashes the calling agent.
- **Bounded** - `k` is clamped to `[1, 20]`; empty queries are rejected.
- **Version-pinned** deps (`requirements.txt`).
## Stack
| Layer | Choice |
|---|---|
| Embeddings | local ONNX `all-MiniLM-L6-v2` (384-dim, CPU, $0) -- **default**. `bge-large-en-v1.5` (1024-dim, 512-token context) available opt-in via `RAG_MCP_EMBEDDER=bge`; see [CUTOVER.md](./CUTOVER.md). |
| Vector store | ChromaDB embedded `PersistentClient` (zero-infra) |
| Server | `mcp` Python SDK 2.x, stdio transport, protocol revision **2026-07-28** |
## Protocol revision
Pinned to `mcp==2.0.0`, the first SDK release implementing MCP protocol revision
**2026-07-28**. The server serves **both eras on the same stdio connection** -- the
client's first frame picks:
| Client opens with | Negotiated revision | Notes |
|---|---|---|
| a per-request `_meta` envelope (or a `server/discover` probe) | `2026-07-28` | stateless per-request envelope; no `initialize` |
| the classic `initialize` handshake | `2025-11-25` | handshake era caps here -- expected, not a downgrade |
`2026-07-28` is **not reachable via the `initialize` handshake**; it is a "modern"
revision reached through `server/discover` or an inline `_meta` version stamp. Era
selection is automatic and per-connection -- there is no server-side flag.
`tests/test_protocol_version.py` asserts both paths end-to-end, so a dependency
rollback that silently drops the server to an older revision fails CI instead of
passing quietly.
## Quick start
```bash
python -m venv .venv && .venv/Scripts/python -m pip install -r requirements.txt
# Ingest a corpus (markdown). Incremental by default: only files whose content
# changed since the last run are re-embedded.
python -m rag_mcp.cli ingest path/to/docs --db ./store.chroma
# Force a rebuild in place (ignore the manifest, re-embed everything)
python -m rag_mcp.cli ingest path/to/docs --db ./store.chroma --full
# One-off query (corpus root = the auth scope)
python -m rag_mcp.cli query "your question" --db ./store.chroma --corpus path/to/docs -k 5
# Same, restricted to one doc_class ("note" or "handoff" -- see "Filtering by
# document class" below)
python -m rag_mcp.cli query "your question" --db ./store.chroma --corpus path/to/docs --doc-class note
# Run as an MCP server (stdio); configure via env first
# RAG_MCP_CORPUS_ROOT, RAG_MCP_DB_PATH, RAG_MCP_COLLECTION, RAG_MCP_EMBEDDER
python run_server.py # operational entrypoint (referenced by mcp.yaml)
python -m rag_mcp # same server, via the packaged console entry point
rag-mcp # after `pip install jaimenbell-rag-mcp` -- console script
```
## Keeping the index fresh (incremental ingest)
Ingest is **incremental by default**. A manifest inside the store dir records a
SHA-256 of each file's decoded text; a run re-embeds only what actually changed,
and prunes what upsert alone never could (chunks of deleted/renamed notes, and
trailing chunks of notes that got shorter).
Measured on a live 2808-file / 26.6 MiB corpus (bge, CPU):
| Run | Cost |
|---|---|
| tick with no changes | **~0.7s** (walk + read + hash everything) |
| full re-embed | ~2h33m (50,109 chunks at ~5.5 chunks/sec) |
That is what makes a frequent schedule affordable: `reingest.bat` is meant to run
**every 15 minutes** instead of once daily at 03:00, which had left a note written
at 03:05 invisible to `search_knowledge` for nearly 24 hours.
The manifest is only trusted when the **run identity** matches -- embedder, embedding
dimension, collection and chunking parameters. Change any of them and every file is
re-embedded, so an embedder swap can never be silently half-applied. A missing,
corrupt, or mismatched manifest, or a manifest against an empty store, all degrade
to a full rebuild; nothing degrades to a wrong skip.
### Snapshot de-duplication
The manifest's skip is a **whole-file** hash, so it cannot see the duplication that
actually hurts retrieval: a daily snapshot series (`fleet-health-2026-07-23.md` and
friends) repeats yesterday's paragraphs verbatim inside a file whose hash still
changed. Measured on the live vault, one `## RED Bots` status line took **five
distinct values across fifteen consecutive files** and crowded a top-10 with
byte-identical copies of itself, burying the document that explained it at rank 16.
Ingest therefore also de-duplicates at **chunk** level, but only within a dated
series and only against the *immediately preceding* snapshot. The first occurrence
is always embedded and keeps its own date as its `source`; later verbatim repeats
are not embedded, and instead extend the survivor's `repeat_dates` metadata, which
`search_knowledge` returns as `snapshot_date` / `also_unchanged_on` /
`snapshots_covered`. So "what did this say on date X" is still answerable -- that is
why the series is de-duplicated rather than excluded. A value that changes and later
returns is kept, because it is a new fact rather than a repeat.
Scope is narrow and stated with the rule in `rag_mcp/snapshots.py`: filename ending
in `-YYYY-MM-DD`, at least 3 such files sharing a directory and stem, byte-identical
under an identical heading. On the live corpus that is 316 of 2,814 files and
collapses 842 of 50,428 chunks (17.5% of series chunks, 1.67% corpus-wide) while
touching zero ordinary notes. Disable with `--no-snapshot-dedupe`.
`--full` rebuilds in place (ignores the manifest, keeps the store); `--clean`
deletes the store first. Both still WRITE a manifest, so the next run is cheap.
`reingest-clean.bat` (weekly) remains a belt-and-braces reset.
## As an MCP server
Register via `mcp.yaml` (validated against mcp-factory's `Manifest` loader). The tool is
`search_knowledge(query, k, doc_class=None)`; it reads the store configured by the
`RAG_MCP_*` env vars.
### Filtering by document class
Every chunk's metadata carries a `doc_class`, set at ingest time. It is `"handoff"` when
the doc's YAML frontmatter has `type: handoff` or a `tags` entry of `handoff`
(case-insensitive), or -- since a doc's frontmatter is optional and the session mirrors
this exists to flag often carry none -- when the file sits directly under a
`handoff_mirror_dir` (default `context/`, configurable via `ingest()`'s
`handoff_mirror_dir=`/`handoff_mirror_basenames=` params or the CLI's
`--handoff-mirror-dir`/`--handoff-mirror-basename` flags) and is named `handoff.md` /
`active.md` / `resume.md`, or matches an anchored "handoff" filename token (e.g.
`handoff-2026-09-03.md`, `morning-dispatch-handoff.md`) -- never a bare substring, so a
title that merely mentions the word (`handoff-skill-redesign-spec.md`) stays `"note"`.
Everything else defaults to `"note"`. Pass `doc_class` to scope a query to one class, e.g.
to keep an agent's own session/handoff bookkeeping out of a knowledge lookup:
```python
from rag_mcp.search import search_knowledge
search_knowledge(
"what did we decide about X", k=5, store=store, corpus_root=root, doc_class="note",
)
```
`doc_class` is a validated, case-sensitive enum -- `"note"` or `"handoff"` (see
`rag_mcp.search.ALLOWED_DOC_CLASSES`) -- or omitted for no filter. A value outside that
set (wrong case, a typo, any other type) returns a structured `invalid_doc_class` error,
same shape as `invalid_query`. A syntactically valid `doc_class` that simply has no
matches in the current store still fails soft to an empty, `ok: true` result.
An incremental ingest run backfills `doc_class` (and any other metadata-schema change) onto
already-embedded, content-unchanged chunks WITHOUT re-embedding them -- see
`ingest.CURRENT_METADATA_VERSION`. A store that predates this feature entirely gets the
correct `doc_class` on every chunk after exactly one incremental run, not a full `--clean`
rebuild.
## Tests
```bash
python -m pytest # 246 passed
```
## Layout
```
rag_mcp/
chunking.py heading-scoped, overlapping markdown chunks
store.py VectorStore (Chroma) + Embedder protocol (MiniLM default + BgeEmbedder opt-in + offline HashEmbedder)
ingest.py idempotent ingest pipeline with source/heading/chunk-index metadata; incremental by default
manifest.py per-file content hashes -> skip unchanged files, prune stale chunks
search.py search_knowledge: cited, auth-scoped, fail-soft, bounded
server.py MCP stdio server exposing search_knowledge
config.py env-driven Config
cli.py ingest + query CLI
__main__.py console entrypoint (`python -m rag_mcp` / `rag-mcp` script); fails loud on missing config
run_server.py operational MCP entrypoint (referenced by mcp.yaml)
mcp.yaml manifest (mcp-factory model)
```
## Commercial support
Maintained by [Jaimen Bell](https://jaimenbell.dev). For production MCP
integrations, custom servers, or agent-reliability work, see
[jaimenbell.dev](https://jaimenbell.dev).
Building your own MCP server? The [MCP Starter Kit](https://jaimenbell.gumroad.com/l/adnojp)
has templates, a build playbook, and packaging war-stories from shipping this one.
<!-- MCP registry ownership marker -->
mcp-name: io.github.jaimenbell/rag-mcp