io.github.ribhav-jain/docsonar
Local document search over your folders: hybrid keyword+semantic, fully offline, read-only
Open source Open in the app JSON README (API)
About
Local document search over your folders: hybrid keyword+semantic, fully offline, read-only
Details
- Kind
- MCP servers
- Topic
- No topic detected
- Publisher
- ribhav-jain
- Origin
- official
- Category
- ferramentas
- Transport
- local
- Version
- 0.1.1
- Last push
- 2026-07-20T18:58:36Z
- Repository state
- ativo
- Language
- Python
- License
- MIT
- Added
- 2026-08-29 04:01:17
- Updated
- 2026-08-29 04:01:17
- Origin id
io.github.ribhav-jain/docsonar
README
# docsonar
<!-- mcp-name: io.github.ribhav-jain/docsonar -->
[](https://pypi.org/project/docsonar/)
[](https://github.com/ribhav-jain/docsonar/actions/workflows/ci.yml)
[](https://pypi.org/project/docsonar/)
[](LICENSE)
**Chat with your folders.** A local document search MCP server — your files never leave your machine.
docsonar indexes folders of documents into a single SQLite database and gives any MCP client (Claude Desktop, Claude Code, and others) hybrid keyword + semantic search over them. Fully offline, zero configuration, read-only by design.
## Quickstart
### Claude Code
Add to your project's `.mcp.json` (or `~/.claude.json` for all projects):
```json
{
"mcpServers": {
"docsonar": {
"command": "uvx",
"args": ["docsonar"]
}
}
}
```
### Claude Desktop
Add to `claude_desktop_config.json` (Settings → Developer → Edit Config):
```json
{
"mcpServers": {
"docsonar": {
"command": "uvx",
"args": ["docsonar"]
}
}
}
```
Running from a source checkout instead: `"command": "uv", "args": ["run", "--directory", "/path/to/docsonar", "docsonar"]`.
Then just talk to it: _"Add my `~/Documents/notes` folder and find everything about quarterly planning."_ The first index downloads the embedding model (~130 MB, one time); keyword search works immediately while that happens.
For HTTP instead of stdio: `uvx docsonar --transport http --port 8365`.
## Tools
| Tool | Purpose |
| --------------------------------------------------- | -------------------------------------------------------------------------------------------------------------- |
| `add_folder(path, include_globs?, exclude_globs?)` | Register a folder; indexing starts in the background |
| `remove_folder(path)` | Unregister and purge its index data |
| `list_folders()` | Registered folders with counts and last index time |
| `search(query, top_k?, folder?, file_type?, mode?)` | Hybrid keyword+semantic search (RRF-fused); ranked passages with file path, heading/page location, and snippet |
| `read_file(path, start_line?, end_line?)` | File text (extracted text for pdf/docx/html); refuses paths outside registered folders |
| `find_similar(path, top_k?)` | Documents most similar in meaning to a given file |
| `reindex(folder?, force?)` | Incremental refresh: new/changed files reindexed, deleted files purged |
| `index_status()` | Index totals, embedding model state, background job progress, failed files |
Supported formats: `txt`, `md`, `pdf` (page-aware, cites `p. 12`), `docx`, `html`.
## Architecture
```mermaid
flowchart LR
Client["MCP client<br/>(Claude Desktop / Code)"] <-->|stdio or HTTP| Server["FastMCP server<br/>8 tools"]
Server --> Sec["path security<br/>(registered folders only)"]
Server --> Search["hybrid search<br/>BM25 + cosine KNN + RRF"]
Server --> Worker["background indexer<br/>(worker thread + queue)"]
Worker --> Parsers["parsers<br/>txt · md · pdf · docx · html"]
Parsers --> Chunker["heading/page-aware chunker<br/>~400 tokens, 60 overlap"]
Chunker --> DB[("SQLite<br/>FTS5 + sqlite-vec")]
Search --> DB
Embed["sentence-transformers<br/>bge-small-en-v1.5 (local)"] --> Worker
Embed --> Search
```
Every chunk stores its location (heading path like `Setup > Windows`, or PDF page range), so search results can cite `report.pdf, p. 12`.
## Design decisions
- **Read-only by design.** There are no write, move, or delete tools, and there never will be. The server only reads files inside folders you explicitly register — with symlink-escape protection and strict path resolution — so the blast radius of a misbehaving client is zero.
- **SQLite as the single store.** FTS5 gives production-grade BM25 keyword search in the standard library, and [sqlite-vec](https://github.com/asg017/sqlite-vec) puts vectors in the same file. One database file, no services to run, trivial to back up or delete.
- **Hybrid search by default.** BM25 and cosine-KNN rankings are fused with reciprocal-rank fusion (k=60). Exact identifiers and rare terms win on the keyword side; paraphrased questions win on the semantic side; RRF needs no score calibration between them. `mode` lets the caller force either side.
- **Local embeddings.** [BAAI/bge-small-en-v1.5](https://huggingface.co/BAAI/bge-small-en-v1.5) (384-dim) — same size class as the classic all-MiniLM-L6-v2 but stronger on retrieval benchmarks. Downloads on first index, runs on CPU, lazy-loaded so server startup stays instant. If the model can't load, everything degrades gracefully to keyword search and tool responses say so.
- **A tool surface built for an LLM caller.** `search` returns enough per hit (path, location, score, snippet) to decide what to read next without another round trip; results carry stable `chunk_id`s; every degradation is reported in-band via `note`/`embedding_note` fields instead of failing. There's no "answer the question" tool on purpose — the calling model does the reasoning; docsonar does retrieval.
- **Incremental by content, not just mtime.** Reindexing checks mtime+size first, then falls back to a SHA-256 content hash — touched-but-identical files are skipped, and deleted files are purged from both the FTS and vector indexes.
- **Scanned PDFs fail loudly.** A PDF with no extractable text is reported as failed with a clear reason rather than silently indexed as empty. OCR is out of scope.
## Benchmarks
Synthetic corpus of 200 markdown files (600 chunks); Intel Core Ultra 5 125U (laptop, CPU-only), Windows 11, Python 3.12. Reproduce with `uv run python scripts/benchmark.py`.
| Operation | Result |
| --------------------------------------- | ------------------------------------- |
| Index, keyword-only | 0.7 s (≈280 files/s) |
| Index, with embeddings | 59 s (≈3.4 files/s — embedding-bound) |
| Incremental reindex, nothing changed | 0.03 s |
| Search, keyword | 8.9 ms median |
| Search, semantic | 35 ms median |
| Search, hybrid | 54 ms median |
| Embedding model load (once per process) | ~21 s |
| Database size | 2.6 MB (1.1 MB keyword-only) |
## Configuration
Zero config needed. To customize, create `config.toml` in the platform config directory (Windows: `%LOCALAPPDATA%\docsonar\`, macOS: `~/Library/Application Support/docsonar/`, Linux: `~/.config/docsonar/`):
```toml
embedding_model = "BAAI/bge-small-en-v1.5" # any sentence-transformers model
chunk_target_tokens = 400
chunk_max_tokens = 512
chunk_min_tokens = 100
chunk_overlap_tokens = 60
max_file_size_mb = 50
extra_ignore_dirs = ["Archive"]
```
CLI flags: `--transport stdio|http`, `--host`, `--port`, `--db-path`, `--config`.
The index database lives in the platform data directory (Windows: `%LOCALAPPDATA%\docsonar\`, macOS: `~/Library/Application Support/docsonar/`, Linux: `~/.local/share/docsonar/`).
## Running from source
Requires [uv](https://docs.astral.sh/uv/) (it installs the right Python automatically):
```sh
# Windows
powershell -ExecutionPolicy ByPass -c "irm https://astral.sh/uv/install.ps1 | iex"
# macOS / Linux
curl -LsSf https://astral.sh/uv/install.sh | sh
git clone https://github.com/ribhav-jain/docsonar.git
cd docsonar
uv sync --dev # creates .venv and installs everything, incl. dev tools
```
### Run the tests
```sh
uv run pytest # full suite (~2 s, no model download needed)
uv run pytest -v # verbose, one line per test
uv run pytest tests/test_search.py # one file
uv run pytest -k incremental # tests matching a keyword
uv run ruff check . # lint
uv run mypy src tests # strict typecheck
uv run python scripts/benchmark.py # perf numbers (downloads the real model)
```
### Run the server
```sh
uv run docsonar # stdio — for MCP clients (Claude Desktop/Code)
uv run docsonar --transport http --port 8365 # HTTP — for manual testing (Postman, curl)
```
VS Code: press F5 — launch configs for the HTTP server and the test suite are in `.vscode/launch.json`.
### Try the tools without an MCP client
With the HTTP server running, the endpoint is `http://127.0.0.1:8365/mcp`. Recent Postman versions can connect directly (**New → MCP Request**, transport HTTP) and show all 8 tools as forms. For raw HTTP, MCP is JSON-RPC over POST with two setup calls, then tool calls. Send every request with headers `Content-Type: application/json` and `Accept: application/json, text/event-stream`:
```jsonc
// 1. initialize — copy the `mcp-session-id` RESPONSE header and send it
// back as a request header on every call below
{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2025-03-26","capabilities":{},"clientInfo":{"name":"postman","version":"1.0"}}}
// 2. initialized notification (expect 202, empty body)
{"jsonrpc":"2.0","method":"notifications/initialized"}
// 3. list the tools
{"jsonrpc":"2.0","id":2,"method":"tools/list"}
```
Then call tools — the repo ships a sample corpus in `examples/sample-docs` to play with (use forward slashes in JSON to avoid escaping; they work on Windows too):
```jsonc
// Register the sample folder (first ever call downloads the embedding model, ~30 s)
{"jsonrpc":"2.0","id":3,"method":"tools/call","params":{"name":"add_folder",
"arguments":{"path":"C:/path/to/docsonar/examples/sample-docs"}}}
// Watch indexing progress until indexing.active is false
{"jsonrpc":"2.0","id":4,"method":"tools/call","params":{"name":"index_status","arguments":{}}}
// Hybrid search — finds the expense-policy PDF, cites its page
{"jsonrpc":"2.0","id":5,"method":"tools/call","params":{"name":"search",
"arguments":{"query":"meal reimbursement limit","top_k":3,"mode":"hybrid"}}}
// Semantic search — no keyword overlap needed
{"jsonrpc":"2.0","id":6,"method":"tools/call","params":{"name":"search",
"arguments":{"query":"how hot should water be for green tea","mode":"semantic"}}}
// Read a hit (works on PDFs too — returns extracted text)
{"jsonrpc":"2.0","id":7,"method":"tools/call","params":{"name":"read_file",
"arguments":{"path":"C:/path/to/docsonar/examples/sample-docs/expense-policy.pdf"}}}
// "More like this"
{"jsonrpc":"2.0","id":8,"method":"tools/call","params":{"name":"find_similar",
"arguments":{"path":"C:/path/to/docsonar/examples/sample-docs/espresso-guide.md","top_k":3}}}
// Refresh after files change on disk (force:true rebuilds everything)
{"jsonrpc":"2.0","id":9,"method":"tools/call","params":{"name":"reindex","arguments":{}}}
// Clean up — unregisters the folder and purges its index data
{"jsonrpc":"2.0","id":10,"method":"tools/call","params":{"name":"remove_folder",
"arguments":{"path":"C:/path/to/docsonar/examples/sample-docs"}}}
```
Security check worth trying: `read_file` with any path outside a registered folder (e.g. `C:/Windows/System32/drivers/etc/hosts`) returns a refusal, not file content.
CI runs lint, typecheck, and the test suite on Linux, macOS, and Windows.
## License
MIT