{
  "markdown": "# understory 🌱\n\n**Memory that grows.**\n\nThe layer beneath your agents: a self-wiring, plain-markdown memory. Every fact your agents learn is filed as a markdown concept, cross-linked into a living knowledge graph, and kept healthy by the agent itself — searchable, diffable, and entirely yours. Runs great on local models.\n\nBundles follow the [Open Knowledge Format (OKF) v0.1 spec](https://github.com/GoogleCloudPlatform/knowledge-catalog/blob/main/okf/SPEC.md) — plain markdown files with YAML frontmatter, readable by humans, diffable in git, portable across tools.\n\n**Three ways in, one agent:**\n\n- **MCP server** — `memory_query` / `memory_add` / `memory_update` / `memory_status` / `memory_maintain` tools over stdio or streamable HTTP. Each call drives an internal LLM agent with the OKF spec in its system prompt.\n- **Web UI** — browse the bundle (tree, concept viewer, update log, conformance badge), see the memory as an Obsidian-style **force-directed graph** (drag/pan/zoom, colored by type, sized by connections, orphans ringed red, click to open), and chat with the same agent to test it. Tool calls render inline so you can watch it work.\n- **Query-path replay** — every agent run (query/mutation/chat) records its traversal (searches → reads → writes) as a compact notation, persisted under `<bundle>/.traces/`. The graph view lists recent runs; selecting one replays the path as numbered directed hops over the graph — visited concepts ringed, search hits dotted, everything else faded.\n- **CLI** — `pnpm agent:query \"...\"` / `pnpm agent:mutate \"...\"` smoke entries.\n\n**Design rule: conformance is enforced in code, not prompts.** The deterministic bundle layer validates frontmatter (`type` required), regenerates `index.md` files, appends `log.md` entries (newest-first, spec §7), and sandboxes all paths to the bundle root. The LLM decides *what* to change; the code guarantees the result is a conformant bundle.\n\n## Quick start (Docker)\n\nNo clone needed — the image is public. Save this as `docker-compose.yml`:\n\n```yaml\nservices:\n  understory:\n    image: ghcr.io/thecodacus/understory:latest\n    ports:\n      - \"3800:3800\"\n    # Lets the container reach a llama.cpp server running on the host via\n    # http://host.docker.internal:8080/v1 (see \"Local llama.cpp\" below).\n    extra_hosts:\n      - \"host.docker.internal:host-gateway\"\n    volumes:\n      # Your memory lives here as plain markdown — a named volume, or point\n      # a bind mount (e.g. ./my-memory:/bundle) at any OKF bundle.\n      - understory-memory:/bundle\n    environment:\n      BUNDLE_ROOT: /bundle\n      LLM_API_BASE_URL: ${LLM_API_BASE_URL}\n      LLM_API_KEY: ${LLM_API_KEY}\n      LLM_API_FORMAT: openai\n      LLM_MODEL: ${LLM_MODEL:-}\n      # Optional fallback\n      LLM_FALLBACK_API_BASE_URL: ${LLM_FALLBACK_API_BASE_URL:-}\n      LLM_FALLBACK_API_KEY: ${LLM_FALLBACK_API_KEY:-}\n      LLM_FALLBACK_API_FORMAT: ${LLM_FALLBACK_API_FORMAT:-openai}\n      LLM_FALLBACK_MODEL: ${LLM_FALLBACK_MODEL:-}\n    restart: unless-stopped\n\nvolumes:\n  understory-memory:\n```\n\n```bash\ndocker compose up -d\n```\n\n### Choosing a provider\n\nThe generic provider system supports any OpenAI-compatible or Anthropic-compatible API.\nSet `LLM_API_BASE_URL` + `LLM_API_KEY` + `LLM_MODEL` and leave `LLM_PROVIDER` unset.\n\n**DeepSeek:**\n```bash\nLLM_API_BASE_URL=https://api.deepseek.com/v1 LLM_API_KEY=sk-... LLM_MODEL=deepseek-chat\n```\n\n**OpenAI:**\n```bash\nLLM_API_BASE_URL=https://api.openai.com/v1 LLM_API_KEY=sk-... LLM_MODEL=gpt-4o\n```\n\n**Anthropic (Claude):**\n```bash\nLLM_API_BASE_URL=https://api.anthropic.com/v1 LLM_API_KEY=sk-ant-... LLM_API_FORMAT=anthropic LLM_MODEL=claude-sonnet-5\n```\n\n**Groq:**\n```bash\nLLM_API_BASE_URL=https://api.groq.com/openai/v1 LLM_API_KEY=gsk_... LLM_MODEL=llama-3.3-70b-versatile\n```\n\n**Local llama.cpp:**\n```bash\nLLM_API_BASE_URL=http://host.docker.internal:8080/v1 LLM_MODEL=\n```\n\n> When understory runs in Docker, `localhost` is the container itself, not the\n> host — so a llama-server on the host is reached at `host.docker.internal`\n> (the compose files above already map it via `extra_hosts`). Running from\n> source on the same box as llama-server, use `http://localhost:8080/v1`.\n\n**Local llama.cpp with DeepSeek fallback:**\n```bash\nLLM_API_BASE_URL=http://host.docker.internal:8080/v1 LLM_MODEL= \\\nLLM_FALLBACK_API_BASE_URL=https://api.deepseek.com/v1 LLM_FALLBACK_API_KEY=sk-... LLM_FALLBACK_MODEL=deepseek-chat\n```\n\nThe old `LLM_PROVIDER` + per-provider key env vars still work (backward-compatible) but are deprecated.\n\nThen:\n\n- **Web UI** → http://localhost:3800 — browse the memory, watch the graph, chat with the agent\n- **MCP endpoint** → `http://localhost:3800/mcp` (streamable HTTP) — register it in any MCP client:\n  ```bash\n  claude mcp add --transport http ustory http://localhost:3800/mcp\n  ```\n- Your agent now has `memory_query` / `memory_add` / `memory_update` / `memory_status` / `memory_maintain`, and gets a seed overview of the memory at every session start.\n\nTeach it something (`memory_add`: \"We deploy on Fridays, never Mondays\"), then open the graph and watch the concept wire itself in. Deploying with Portainer? Use [docker-compose.portainer.yml](docker-compose.portainer.yml) as a repository stack.\n\n## Stack\n\npnpm monorepo:\n\n| Package | What |\n|---|---|\n| `packages/core` | OKF bundle layer (zero LLM) + agent (Vercel AI SDK tool loop: search/read/list/write/patch/delete) + provider registry |\n| `packages/server` | Express: MCP streamable-HTTP at `/mcp`, stdio bin, REST browse API at `/api/*`, streaming chat at `/api/chat`, serves the web build |\n| `packages/web` | Vite + React + TS + Tailwind: bundle browser + agent chat (`useChat`) |\n\nProviders are configured through `LLM_API_BASE_URL`, `LLM_API_KEY`, `LLM_API_FORMAT` (`openai` or `anthropic`), and `LLM_MODEL`. Any OpenAI-compatible endpoint (DeepSeek, OpenAI, Groq, OpenRouter, llama.cpp, etc.) works with `LLM_API_FORMAT=openai`; Anthropic-compatible endpoints use `LLM_API_FORMAT=anthropic`. Optional fallback uses the matching `LLM_FALLBACK_*` variables.\n\n### llama.cpp\n\n```bash\n# on the inference box — --jinja enables OpenAI-style tool calling\nllama-server -m model.gguf --jinja --host 0.0.0.0 --port 8080\n\n# here — no model id needed, it's discovered for llama-server-like local endpoints\nLLM_API_BASE_URL=http://inference-box:8080/v1 LLM_API_FORMAT=openai LLM_MODEL= \\\nBUNDLE_ROOT=./sample-bundle node packages/server/dist/index.js\n```\n\nWorks behind llama-swap too: discovery prefers the currently **loaded** model so a query doesn't trigger a multi-minute model swap. Pin a specific model with `LLM_MODEL=`.\n\n## From source\n\n```bash\npnpm install\npnpm build\ncp .env.example .env   # add your API key\n\nBUNDLE_ROOT=./sample-bundle \\\nLLM_API_BASE_URL=https://api.deepseek.com/v1 \\\nLLM_API_KEY=sk-... \\\nLLM_API_FORMAT=openai \\\nLLM_MODEL=deepseek-chat \\\nnode packages/server/dist/index.js\n# → http://localhost:3800  (web UI + /api + /mcp)\n```\n\nOr build the container yourself: `docker compose up --build` (the repo's [docker-compose.yml](docker-compose.yml) builds from source and mounts `./sample-bundle`).\n\nDev mode (server on :3800, Vite HMR on :5180 with proxy):\n\n```bash\nBUNDLE_ROOT=./sample-bundle pnpm --filter @understory/server dev\npnpm --filter @understory/web dev\n```\n\n## MCP registration (Claude Code / Desktop)\n\n```bash\nclaude mcp add ustory \\\n  -e BUNDLE_ROOT=/path/to/your/bundle \\\n  -e LLM_API_BASE_URL=https://api.deepseek.com/v1 \\\n  -e LLM_API_KEY=sk-... \\\n  -e LLM_API_FORMAT=openai \\\n  -e LLM_MODEL=deepseek-chat \\\n  -- node /path/to/understory/packages/server/dist/mcp/stdio.js\n```\n\nOr point an HTTP MCP client at `http://host:3800/mcp`.\n\n### Auth\n\nBy default the server is open — fine on localhost or a trusted LAN. Before exposing it anywhere else, set `AUTH_TOKEN`:\n\n```bash\nAUTH_TOKEN=$(openssl rand -hex 24)\n```\n\nWith it set, `/mcp` and `/api` require `Authorization: Bearer <token>` (the web UI stays reachable and prompts for the token). Register authenticated MCP clients with a header:\n\n```bash\nclaude mcp add --transport http ustory http://host:3800/mcp \\\n  --header \"Authorization: Bearer <token>\"\n```\n\nThe stdio transport needs no token — it's a local process spawned by the client.\n\n### Seed memory\n\nA client LLM that only sees four bare tool names never gets the instinct to check memory. So at **session start** the server injects a compact overview of what the knowledge base contains (directories, concepts with types + descriptions, recent activity) through both channels that reach the model:\n\n1. the MCP initialize **`instructions`** field (clients like Claude put it in the system prompt), and\n2. the **`memory_query` tool description** — the universal fallback every tool-calling client loads.\n\nThe seed regenerates fresh for every new session. After `memory_add` / `memory_update` in a long-lived (stdio) session, the tool description refreshes via `tools/list_changed`, so the session sees its own writes. Out-of-band edits (hand edits, other clients) are picked up on the next session.\n\n### Graph health & maintenance\n\nMemory is a graph, not a pile of notes, and graphs rot: concepts go **orphaned** (nothing links to them) and links go **broken**. Two mechanisms keep it healthy:\n\n- **Write-time linking** — new knowledge either enriches the concept it belongs to (an attribute of an existing entity is patched in, not filed separately) or, when it's a distinct entity, is created *and* back-linked from related concepts. Contradictions are superseded in place, never left standing alongside the old value.\n- **`memory_maintain`** — a deterministic lint (orphans + broken links, surfaced in `memory_status` under `graph`) drives an internal agent to wire orphans into related concepts and fix dangling links. Run it periodically to counter drift; it's a no-op when the graph is already healthy.\n\nThis design mirrors the pattern in Karpathy's [LLM Wiki](https://gist.github.com/karpathy/442a6bf555914893e9891c11519de94f) (index.md + log.md, create-vs-enrich, lint for orphans). Deferred from that pattern until scale warrants: an explicit page-type schema, and hybrid FTS5+embedding search (the naive scan in `search.ts` is fine into the low thousands of concepts).\n\n## Tests\n\n```bash\npnpm test                                  # core: 18 tests (spec §5/§6/§7/§9, sandbox, search, concurrency)\npnpm --filter @understory/server exec tsx scripts/mcp-smoke.mts   # MCP stdio round-trip (needs SMOKE_BUNDLE + an API key)\n```\n\n## Environment\n\nSee [.env.example](.env.example). `BUNDLE_ROOT` is required; `GIT_AUTOCOMMIT=true` commits every mutation.\n",
  "bytes": 10591,
  "sha": "bf319e4787348a8acce6058abaaf3addb9628cc2edb5bef97c8ccda4310a20c8",
  "repo_slug": "thecodacus/understory",
  "fonte": "repo",
  "truncated": false,
  "api": "https://agentalog.com/api/listings/okf_thecodacus_understory_sample_bundle_inde_e3497965/readme"
}