Docs Ask (local docs RAG)
Local RAG MCP for markdown documentation. Retrieval only; host synthesizes answers.
Open source Open in the app JSON README (API)
About
Local RAG MCP for markdown documentation. Retrieval only; host synthesizes answers.
Details
- Kind
- MCP servers
- Topic
- AI, RAG & memory
- Publisher
- alyiox
- Origin
- official
- Category
- ferramentas
- Transport
- local
- Version
- 0.3.0
- Last push
- 2026-09-03T08:59:50Z
- Repository state
- ativo
- Language
- Python
- License
- MIT
- Added
- 2026-08-29 03:02:24
- Updated
- 2026-09-03 10:00:16
- Origin id
io.github.alyiox/mcp-docs-ask
README
# mcp-docs-ask
[](https://github.com/alyiox/mcp-docs-ask/actions/workflows/ci.yml)
[](https://pypi.org/project/mcp-docs-ask/)
[](https://www.python.org/downloads/)
[](LICENSE)
<!-- mcp-name: io.github.alyiox/mcp-docs-ask -->
Local RAG MCP for documentation. Point `source` at any markdown repository
(local path or git URL).
The server does **retrieval only** (no answer LLM). `ask_docs` returns grounded
passages and citations; the MCP host (Cursor / Claude) synthesizes the answer.
## Features
- `ask_docs` retrieval with configurable path-based layer filters
- `list_docs` discovery for configured docs collections and layer filters
- `reindex` rebuilds the local vector index; for git URL sources it also fetches updates
## Requirements
- Python 3.13+
- [`uv`](https://docs.astral.sh/uv/)
- `git` on PATH (only if `source` is a git URL)
- Git credentials on the machine when `source` is a **private** git URL
(`gh auth login`, HTTPS credential helper, or SSH). No tokens in config.
- First run downloads the embedding model weights once (sentence-transformers)
## Quick start
```bash
git clone git@github.com:alyiox/mcp-docs-ask.git
cd mcp-docs-ask
uv sync
mkdir -p ~/.config/mcp-docs-ask
cp config.example.json ~/.config/mcp-docs-ask/config.json
# Prefer a local checkout while developing:
# set docs.<id>.source to your docs repo path
npx -y @modelcontextprotocol/inspector uv run mcp-docs-ask
```
## Configuration
Config path: `~/.config/mcp-docs-ask/config.json`
> **Windows:** `%USERPROFILE%\.config\mcp-docs-ask\config.json`
```json
{
"docs": {
"product": {
"source": "https://github.com/example/docs.git",
"desc": "Product guides and API reference",
"ref": "main",
"include": ["**/*.md"],
"exclude": ["archive/**"],
"layers": {
"guides": {
"desc": "How-to and onboarding guides",
"include": ["docs/guides/**"]
},
"api": {
"desc": "HTTP API reference",
"include": ["docs/api/**"]
}
},
"embedding_model": "sentence-transformers/all-MiniLM-L6-v2"
},
"team-notes": {
"source": "/path/to/docs",
"desc": "Internal team notes (local path; ref unused)",
"include": ["**/*.md"],
"exclude": ["archive/**"]
}
},
"default": {
"docs": "product",
"embedding_model": "sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2",
"top_k": 8,
"chunk_max_chars": 1500
}
}
```
`product` is a git URL (`ref` applies). `team-notes` is a filesystem path (`ref` unused).
Optional `desc` on each docs collection and layer helps agents pick the right target.
`embedding_model`, `top_k`, and `chunk_max_chars` resolve as:
`docs.<id>.X` → `default.X` → built-in. Omit per-docs keys to inherit.
**Embedding model recommendation**
Any Hugging Face id loadable by `sentence-transformers` works. Pick by language mix:
| Docs / queries | Recommended `embedding_model` |
|---|---|
| **English-only** (built-in when omitted) | `sentence-transformers/all-MiniLM-L6-v2` |
| **Chinese-only** | `BAAI/bge-small-zh-v1.5` |
| **Multilingual** (~50 langs) | `sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2` |
Changing `embedding_model` requires a `reindex` (the on-disk index stores the model name).
| Field | Description |
|---|---|
| `docs.<id>.source` | Docs **repo root**: local path or git URL |
| `docs.<id>.desc` | Short description for discovery (`list_docs`) |
| `docs.<id>.ref` | Branch / tag / SHA for git URL sources only (default `main`; ignored for local paths) |
| `docs.<id>.include` | Globs relative to repo root (default `**/*.md`) |
| `docs.<id>.exclude` | Globs to skip |
| `docs.<id>.layers.<name>.include` | Path globs for that layer (first match wins) |
| `docs.<id>.layers.<name>.desc` | Short layer description for discovery |
| `docs.<id>.embedding_model` | Optional override (see recommendation above) |
| `docs.<id>.top_k` | Optional override for default retrieval count |
| `docs.<id>.chunk_max_chars` | Optional override for max body chars per heading chunk |
| `default.docs` | Default docs collection id |
| `default.embedding_model` | Default sentence-transformers model id |
| `default.top_k` | Default retrieval count |
| `default.chunk_max_chars` | Default max body chars per heading chunk |
**Layers** partition indexed files by path glob. First match wins. Names are
case-insensitive; `all` is reserved (cannot be configured as a layer name).
| `ask_docs` `layer` | Meaning |
|---|---|
| `all` (default) | Every indexed chunk (named layers and paths outside them) |
| `<named>` | Only chunks whose path matched that named layer’s `include` globs |
Paths that match no named-layer glob are still indexed and only appear under
`layer=all`. Omit `layers` (or set `"layers": {}`) for flat repos — use
`layer=all`.
Cache layout:
- Repos (git URL): `~/.cache/mcp-docs-ask/repos/<docs-id>/`
- Indexes: `~/.cache/mcp-docs-ask/indexes/<docs-id>/`
## Tools
| Tool | Description |
|---|---|
| `list_docs` | List configured docs collections, layer filters, and index state |
| `ask_docs` | Retrieve grounded passages + citations (`layer`: `all` or a named layer) |
| `reindex` | Sync git source (if URL) and rebuild the vector index |
`list_docs` returns a `default` block with the same keys as the config `default`
block (`docs`, `embedding_model`, `top_k`, `chunk_max_chars`), plus a `docs` list
where each entry carries its resolved values, a `default` flag, and an `index`
block (`null` when the collection has never been indexed). Valid `layer` values
are `all` plus the named layer ids — see **Layers** above.
### Index block
`list_docs` and `reindex` return the same `index` keys: `origin`, `root`, `rev`,
`files`, `chunks`, `layers`.
`origin` mirrors the configured source: `file` for a filesystem path, `git` for
a URL the server clones into `~/.cache/mcp-docs-ask/repos/<docs-id>/` and
fetches on `reindex`. `root` is where the files actually are — `null` only when
a built index outlived its source directory. `rev` is the checkout HEAD when
there is one, so a `file` source that is itself a git clone still reports one;
its working tree may hold uncommitted edits, so `rev` labels the checkout, not
the exact indexed content.
`ask_docs` carries only the two answer-scoped keys, `root` and `rev`: the
checkout that produced the passages, and the revision they came from.
### Reading a full source file
`citations[].path` is repo-relative and stable; `answer_context` holds the
passage text once, keyed by the `[n]` markers that match `citations[].n`. To read
a whole source file, join `index.root` from the same `ask_docs` response with a
citation path:
```
/home/you/docs-repo + product/features/budget.md
```
Take `root` from the response that produced the citations rather than an earlier
`reindex` — `ask_docs` rebuilds a stale index itself, so its `rev` is the one
that matches the passages in hand.
## MCP host examples
The examples below launch the server with `uvx`, which installs the package on first
use. Run it once in a terminal beforehand so your host does not block on that install:
```bash
$ uvx mcp-docs-ask
Installed 84 packages in 275ms
```
The server then starts on stdio and waits for input — press Ctrl-C once you see the
install line. Embedding model weights are fetched separately, on the first `ask_docs`
or `reindex` call.
> **Linux (including WSL, containers, and CI):** the PyPI `torch` wheel for Linux is
> the CUDA build. It pulls ~15 `nvidia-*` packages whether or not the machine has an
> NVIDIA GPU — about 2.7 GB of wheels and ~4 GB on disk. Windows and macOS resolve to
> a CPU-only wheel (~1 GB) and never download CUDA. Pre-warming matters most here:
> expect the first `uvx` run to take minutes, not milliseconds.
### Cursor
Add to `.cursor/mcp.json`:
```json
{
"mcpServers": {
"docs-ask": {
"command": "uvx",
"args": ["mcp-docs-ask"]
}
}
}
```
### Claude Code
Add to your Claude Code MCP config:
```json
{
"mcpServers": {
"docs-ask": {
"command": "uvx",
"args": ["mcp-docs-ask"]
}
}
}
```
### Codex
```toml
[mcp_servers.docs-ask]
command = "uvx"
args = ["mcp-docs-ask"]
```
### OpenCode
```json
{
"$schema": "https://opencode.ai/config.json",
"mcp": {
"docs-ask": {
"type": "local",
"enabled": true,
"command": ["uvx", "mcp-docs-ask"]
}
}
}
```
### GitHub Copilot
```json
{
"inputs": [],
"servers": {
"docs-ask": {
"type": "stdio",
"command": "uvx",
"args": ["mcp-docs-ask"]
}
}
}
```
## Development
```bash
uv sync
uv run ruff check src/ tests/
uv run ruff format --check src/ tests/
uv run pyright
uv run pytest
```
## Notes
- **Local path:** `ask_docs` rebuilds the index automatically when file mtimes/sizes
change (fingerprint check). You do not need `reindex` after editing local docs.
- **Git URL:** `ask_docs` never fetches. Call `reindex` to `git fetch` the configured
`ref` and rebuild.
- Changing `embedding_model` invalidates the on-disk index (rebuild on next use /
`reindex`).