{
  "markdown": "<div align=\"center\">\n\n<img src=\"https://raw.githubusercontent.com/true-alter/mcp-ollama/main/docs/alter-mark.svg\" alt=\"\" height=\"96\">\n\n# ~alter mcp-ollama\n\n**Hands the work that shouldn't leave your machine to the model already sitting on it**\n\n[![~alter](https://img.shields.io/badge/~alter-identity%20infrastructure-C9A84C?style=flat-square)](https://truealter.com)\n[![MCP](https://img.shields.io/badge/MCP-stdio-555?style=flat-square)](https://modelcontextprotocol.io)\n[![Node](https://img.shields.io/badge/node-%E2%89%A518-555?style=flat-square)](#install)\n[![Licence](https://img.shields.io/badge/licence-Apache--2.0-555?style=flat-square)](./LICENSE)\n\n[What it does](#what-is-mcp-ollama) · [Install](#install) · [The tools](#the-tools) · [Why this sits under ~alter](#why-this-sits-under-alter)\n\n</div>\n\n## What is mcp-ollama?\n\nAn MCP server that hands work to [Ollama](https://ollama.com) on the same\nmachine and passes the answer back. Ten tools over stdio. Your client calls one\nof them, Ollama does the generating on your own GPU, and nothing is charged to\nan API account.\n\nMost of a working session is mechanical. Docstrings, commit messages, PR\ndescriptions, changelog entries, classification and tagging, summarising a long\nfile, converting one format into another, and having a small vision model look\nat a screenshot and report what's on screen. None of that needs a frontier\nmodel, and most of it gets one anyway, because that's what your client already\nhas an API key for.\n\nHand it a staged diff and a commit message comes back. Hand it a chunk of\nsource and you get a docstring, a test stub or a set of type annotations. All\nten tools, and what each one takes, are in [the tools](#the-tools).\n\nThe orchestrator decides what gets routed here. This server makes no judgement\nabout what belongs local, and it doesn't stream, queue or cache. It keeps\nnothing between calls beyond a random identifier for the process it's running\nin.\n\nIt depends on two things and ships neither. Node 18 or newer runs it, and a\nrunning Ollama with at least one model pulled does the actual generating. There\nare no weights in this repository and no download of weights at install time.\nThe default model is `hermes3:8b`, which you can override per call or per\nenvironment.\n\n## Why this sits under ~alter\n\nYou already decided that some work stays here. That's why there's a model on\nthis disk, pulled once and left there, instead of an API key doing the same job\nfrom a datacentre you'll never see. ~alter starts from the same decision and\npoints it at the other thing that leaves your machine constantly, which is the\nrecord of who you are.\n\nNothing on this machine answers that well today. Your account is a password at a\nlogin screen and a token in a config file, and each of those checks one moment\nand then stops looking. The years of work that actually say who you are sit\noutside anything either of them can see. An agent commits under your name now,\nin your editor, and when somebody asks later who allowed that, there's no\nanswer written down anywhere.\n\n~alter answers that with a handle. `~yourname`, one of them, carrying a record\nbuilt out of work you actually did. A client speaking the protocol picks it up\nwithout being configured again, much the way this server picks up a model you\npulled once and left on the disk.\n\nYou bought hardware to stop paying per call. On that record the per-call money\nruns the other way. The bare fact that you exist is free for anyone to check,\ndepth past that is charged to whoever wants it, and you're the one who decides\nwhich readers reach which depth.\n\nNone of it is needed here. `mcp-ollama` has no login and asks for no handle, and\nit behaves the same whether you hold one or not. There's one small wire between\nthem and it's worth naming plainly. If the ~alter runtime happens to be\ninstalled on this machine, every local generation appends a single line to a\nfile in your own home directory, recording that the work ran here at no API\ncost. That line never goes anywhere, and if the runtime isn't installed no line\ngets written at all.\n\n<details><summary><b>I want to know more</b></summary><br><p>Your friends do not know you from a login. Neither does your family, or the people you work with, or your sports team. They know who you are from how you have shown up, over years. You may look and sound nothing like you did ten years ago and it is still you.</p><p>Software still asks the narrow question. A password at the login screen. A token in the app. Each one checks whether this is the right person, right now, at this exact spot, and then looks away. Everything in between is invisible to it, and that is almost all of your life. The AI tools made it worse, because one of them writes in your name now, and when somebody asks who allowed it, there is no answer anywhere on this machine.</p><p><b>One name, and the record under it is yours.</b> <code>~yourname</code> works at every tool that speaks the protocol, so nothing is set up twice and there is no key to paste, at any point. What is in the record was read from what you did, not typed into a box by you or guessed at by a model. You can cut off access to any of it with one command. No form, no review.</p><p><b>Being read pays you.</b> Anyone can check for free that your handle exists. Reading further, past the fact of you, costs the reader money, that money goes to you, and you decide who gets to read what. Collectives hold a record on the same terms, because it is one record at every scale and the same form holds you, your team, and your country. A union, a club, a trade body, or a group of people who have never met can vouch for their own members and be read the same way you are.</p><p><b>Threads are how you get somewhere.</b> A thread is something you are working towards. You take one on yourself, or ~Alter plants one from what it can already see, or a collective confers one on you. What gets read is the trail you leave while you are doing it, never a certificate at the end.</p><p>All of that is the short version of an argument set out properly, in eight chapters at <a href=\"https://truealter.com/origin\">truealter.com/origin</a>. Its seventy-six predictions were each written so the data could come back and say no. Five of them did, and those five are still published alongside the rest.</p></details>\n\n<details><summary><b>I want to check it myself</b></summary>\n\nNone of it needs taking on trust. Hand this to the AI you already use and\nlet it take you through in your own words.\n\n```text\nGo to truealter.com, find the genesis and the origin papers, and read all of\nthem. Then start a discussion with me in my own language, no jargon.\n\nAsk me one question at a time and wait for my answer. Follow what I am curious\nabout, not the order the site is written in.\n\nDo not only read the pages. Connect to their interface directly and call\nget_started. That tells you where I am and what comes next, and we work the\nrest of this from there.\n\nIf I have a ~handle I will give it to you. If I do not and I decide I want\none, set it up with me here.\n\nWhat I want to understand is what ~Alter is, how it applies to me, and what it\nmeans for everyone else. Find me somewhere in it, and not the version written\nfor people who work in technology. Find the life nearest mine and walk me\nthrough it, including what it would cost me.\n\nTry to prove them wrong. Go and find where they show their working, and if\nsomething is unclear or does not hold up, say so.\n```\n\n</details>\n\nThe ~alter command line is where a person meets ~alter, and this isn't that.\n`mcp-ollama` has no command of its own worth typing and never asks you to log\nin to anything. It's a process your MCP client starts and stops, and it\nhappens to be maintained by the same people.\n\n## Install\n\n```bash\nnpm install -g @truealter/mcp-ollama\n```\n\nThat puts one command on your PATH, `mcp-ollama`, which is the process your\nclient launches. The package ships the build already done, so there is no\ncompile step and no toolchain to have installed first. CI builds it against\nNode 18, 20 and 22 before publish, so anything in that range is known to work\nat runtime.\n\nIf you would rather install nothing at all, `npx -y @truealter/mcp-ollama`\nfetches it on first use and runs the same process. That is the form used in\nthe client configuration below.\n\nThe scoped name is the one to type. The unscoped `mcp-ollama` on the public\nregistry belongs to an unrelated publisher and is not this package.\n\nNothing is installed as a service and nothing runs in the background. Your\nclient starts the process when it needs it and stops it when it is done.\n\n## Routing your first job\n\n### 1. Pull the model it reaches for by default\n\n```bash\nollama pull hermes3:8b\n```\n\nThat's the default this server reaches for when a tool call doesn't name a\nmodel. It's quick and it's honest at classification, tagging and short\ngenerations. Heavier models are worth having for code work, and\n[choosing a model](#choosing-a-model) covers when to bother.\n\n### 2. Register the server with your client\n\n```bash\nclaude mcp add --transport stdio ollama -- npx -y @truealter/mcp-ollama\n```\n\nIf you installed globally, `-- mcp-ollama` works just as well and skips the\nfetch. Cursor, Cline and anything else MCP-aware take the same shape in their\nown config. The client launches the process; you never run it by hand except\nto debug, and if you do, it sits there waiting on stdin, which is correct\nrather than broken.\n\n### 3. Ask your client what is on the host\n\n```text\nList the models on the local Ollama host.\n```\n\nSay that to your client in whatever words you like. It resolves to\n`local_models`, which reads Ollama's tag list and reports each model's size,\nparameter count, quantisation and family. If what you just pulled comes back,\nthe wire is good end to end.\n\n### 4. Hand it a diff and ask for a commit message\n\n```bash\ngit diff --staged\n```\n\nHand that output to your client and ask for a commit message. It routes to\n`local_diff` with `commit-message`, which prompts for imperative mood, a subject\nunder 72 characters and a body explaining why rather than what. Nothing about\nthat needed a frontier model, and now it doesn't use one.\n\n## The tools\n\n| Tool | What it does |\n|---|---|\n| `local_generate` | Free-form generation with your own system prompt, temperature and token ceiling |\n| `local_summarize` | Summarise bulk text as bullets, a paragraph or one line, optionally focused on a theme |\n| `local_analyze` | Structure pulled out of text, classification, entities or tags, in an output shape you name |\n| `local_draft` | Formulaic prose against a convention you supply |\n| `local_code` | `docstring`, `test`, `explain`, `review`, `types`, `comments` or `refactor-suggest` over a chunk of source |\n| `local_diff` | `commit-message`, `pr-description`, `changelog`, `summary` or `impact` from a diff |\n| `local_transform` | Mechanical pattern transforms, format conversions, renames and syntax migrations |\n| `local_models` | What's on this Ollama host, with sizes and quantisation |\n| `local_pull` | Pull a model onto this host by name, untagged names only |\n| `local_vision` | Have a vision model look at screenshots and report `see`, `emptystate` or `legibility` |\n\nTen of them, and the full schemas come over MCP introspection, so any MCP-aware\nclient enumerates them without being told.\n\nTwo take a `max_tokens` argument. `local_generate` defaults to 2048 and\n`local_summarize` to 1024. The rest set their own ceiling in code, 4096 for\n`local_code` and `local_transform`, 2048 for `local_analyze` and most of\n`local_diff`, 512 for a commit message, 1024 for `local_draft`. If output comes\nback cut short on one of those, split the input rather than hunting for a\nparameter that isn't there. Temperature is exposed on `local_generate` only.\n\n`local_vision` is the odd one and worth a note. It reads pixels and reports what\nis on screen, whether the main content area holds real data or an error, which\nregions exist, what text is clipped or unreadable. It deliberately doesn't rank\nseverity or approve anything, because a small vision model reads a render well\nand judges it badly. Feed it near full resolution, because below about 1280px\nwide it starts inventing data that isn't there.\n\n## Choosing a model\n\n| Variable | Default | What it does |\n|---|---|---|\n| `OLLAMA_HOST` | `http://localhost:11434` | Where Ollama is listening. Loopback only unless you override the gate below |\n| `OLLAMA_MODEL` | `hermes3:8b` | Model used when a tool call doesn't name one |\n| `OLLAMA_VISION_MODEL` | `qwen2.5vl:7b` | Model `local_vision` uses when a call doesn't name one |\n| `MCP_OLLAMA_ALLOW_REMOTE` | unset | Set to `1` to permit a non-loopback `OLLAMA_HOST` |\n\nAny call can name its own `model` and the environment default only applies when\nit doesn't, so one server handles a mixed workload without being reconfigured.\n\n| Workload | Try | Why |\n|---|---|---|\n| Classification, tagging, one-liners | `hermes3:8b` | Fastest round trip, cheap to keep resident |\n| Commit messages, changelogs, summaries | `qwen2.5-14b-instruct` | Better prose, still comfortable on a 16GB card |\n| Code review, docstrings, tests | `qwen2.5-coder:32b` | Code-specialised, worth the extra VRAM |\n| Looking at a render | `qwen2.5vl:7b` | The vision default, and small enough to stay on the GPU |\n\nRun `local_models` at the start of a session on a host you don't know.\n\n<details><summary><h3>Running it in Docker</h3></summary>\n\nNo image is published anywhere, so every path below starts with a build from\nthis repository.\n\n```bash\ndocker build -t mcp-ollama .\n```\n\nThe `Dockerfile` builds on `node:20-alpine` and already sets `OLLAMA_HOST` to\n`http://host.docker.internal:11434`, so the container reaches Ollama on the host\nrather than looking for it inside itself. That address is not loopback from the\nserver's point of view, so the loopback gate refuses it and the process exits at\nstartup unless `MCP_OLLAMA_ALLOW_REMOTE=1` is set as well. The image does not\nset that one, which is why every command here does.\n\nThe image also sets `OLLAMA_MODEL` to `hermes3:8b`. Add `-e OLLAMA_MODEL=...` to\nroute to a different default.\n\n**Docker, on macOS and Windows**\n\n```bash\ndocker run -i --rm -e MCP_OLLAMA_ALLOW_REMOTE=1 mcp-ollama\n```\n\n**Docker, on Linux**\n\n`host.docker.internal` does not resolve there by default, so map it to the\nbridge gateway.\n\n```bash\ndocker run -i --rm \\\n  --add-host=host.docker.internal:host-gateway \\\n  -e MCP_OLLAMA_ALLOW_REMOTE=1 \\\n  mcp-ollama\n```\n\n**Docker Compose**\n\n`docker-compose.yml` ships in this repository, so there is nothing to write. It\ncarries an `extra_hosts` mapping that makes the same file work on Linux as well\nas Docker Desktop.\n\nThis server speaks MCP over stdin and stdout, so it needs a client on the other\nend of the pipe. `docker compose up` starts it with nothing attached and it sits\nthere doing nothing. Use `run`, with `-T` so Compose leaves the pipe alone.\n\n```bash\ndocker compose run --rm -T mcp-ollama\n```\n\n**Pointing a client at it**\n\nAn MCP client launches the server itself, so hand it the whole command rather\nthan a container that is already running.\n\n```json\n{\n  \"mcpServers\": {\n    \"ollama\": {\n      \"command\": \"docker\",\n      \"args\": [\"compose\", \"-f\", \"/path/to/docker-compose.yml\", \"run\", \"--rm\", \"-T\", \"mcp-ollama\"]\n    }\n  }\n}\n```\n\n</details>\n\n<details><summary><h3>When something doesn't work</h3></summary>\n\n**`Ollama error 404` on a tool call**\n\nThat model isn't pulled. Run `ollama pull <name>` from a shell. `local_pull`\nhandles untagged names only, because its validator rejects the colon in a tag\nlike `hermes3:8b`.\n\n**`fetch failed`, or connection refused**\n\nOllama isn't running, or `OLLAMA_HOST` points at the wrong place. Check with\n`curl $OLLAMA_HOST/api/tags`. Inside a container, `localhost` is the container\nitself.\n\n**`OLLAMA_HOST must be loopback`, and the process dies immediately**\n\nThat's the gate doing its job. Point it back at `localhost`, or set\n`MCP_OLLAMA_ALLOW_REMOTE=1` if you genuinely meant a remote host.\n\n**Calls feel slow**\n\nA cold model has to load first, and everything after that in the same Ollama\nprocess is much faster. If the model is larger than your VRAM, Ollama spills\nto CPU, and `ollama ps` will tell you so.\n\n**Vision calls balloon memory or crawl**\n\n`local_vision` caps context at 8192 and holds the model for 30 minutes on\npurpose. A 32K context plus one image pushes past 23GB on a 7900-class card\nand spills to CPU, which is where the cap came from.\n\n**Output stops early**\n\nSee the token ceilings under [the tools](#the-tools). Most tools don't take\n`max_tokens`.\n\n</details>\n\n<details><summary><h3>What this server does and doesn't do on your machine</h3></summary>\n\nIt makes no network call other than to the configured `OLLAMA_HOST`, and by\ndefault that host has to be `localhost`, `127.0.0.1` or `::1`. A non-loopback\nvalue throws at startup rather than quietly sending your prompts somewhere else,\nand getting past that takes a deliberate `MCP_OLLAMA_ALLOW_REMOTE=1`. Point it\nat a remote Ollama on purpose and that endpoint's posture becomes yours.\n\nThere's no telemetry, no analytics, no auto-update check and no model weights in\nthe package. Tool inputs go to Ollama's HTTP API as given and the response comes\nstraight back. Model names passed to `local_pull` are validated against\n`^[a-z0-9][a-z0-9._/-]{0,127}$` before they reach the registry endpoint, so a\ncaller-supplied string can't wander off that path.\n\nOne local write is worth knowing about. If\n`~/.local/share/alter-runtime/lib/substrate-emit.sh` exists, each generation\nappends a row to `~/.local/share/alter-runtime/token-burn.jsonl` recording the\ntool, the model and the token counts. It's fire and forget, it fails silently, it\nnever touches the tool result, and if that helper isn't installed nothing is\nwritten.\n\nTo report a security issue, see [SECURITY.md](./SECURITY.md).\n\n</details>\n\n<details><summary><h3>The protocols underneath it</h3></summary>\n\nThe record formats are open Internet-Drafts, so somebody else's implementation reads and writes the same records this one does without asking us. These are the drafts those formats are specified by. This server does not\nimplement them itself; it runs beside the components that do.\n\n| Draft | What it specifies |\n|---|---|\n| [`compute-location-gate`](https://datatracker.ietf.org/doc/draft-morrison-compute-location-gate/) | Negotiating where an identity inference computes, decided by the provenance class of the signal, before any inference runs. |\n| [`mcp-dns-discovery`](https://datatracker.ietf.org/doc/draft-morrison-mcp-dns-discovery/) | The DNS records that publish a `~handle`, the server that answers for it, and the signed envelope bound to it. |\n\nEighteen drafts make up the whole stack. The rest are on the [IETF datatracker](https://datatracker.ietf.org/doc/search/?name=draft-morrison&activedrafts=on).\n\n</details>\n\n<details><summary><h3>The rest of it</h3></summary>\n\nOne identity rail, several ways in.\n\n| Name | What it is |\n|---|---|\n| **[`@truealter/cli`](https://www.npmjs.com/package/@truealter/cli)** | The command line, and the front door for a person. |\n| **[homebrew-tap](https://github.com/true-alter/homebrew-tap)** | That command line, packaged for macOS and Linux. |\n| **[runtime](https://github.com/true-alter/runtime)** | The daemon that keeps your `~handle` known on your own machine. |\n| **[`@truealter/sdk`](https://www.npmjs.com/package/@truealter/sdk)** | Reading identity from your own code. |\n| **[obsidian](https://github.com/true-alter/obsidian)** | ~Alter inside an Obsidian vault, on-device. |\n| **mcp-ollama** | Local models, for work that should stay on the machine it runs on. **You are here.** |\n\nDocumentation is at [truealter.com/docs](https://truealter.com/docs).\n\nBug reports and small patches are welcome, see\n[CONTRIBUTING.md](./CONTRIBUTING.md). A report is most useful with the tool you\ncalled, the client you called it from, the model you routed to, the full error,\nand your Node and Ollama versions. For a larger design change, open an issue\nfirst so we can agree the scope before you spend time on it.\n\n`mcp-ollama` is small and stays that way. Routing work to a local Ollama process\nis the whole brief.\n\nApache 2.0. See [LICENSE](./LICENSE) for the full text. Copyright 2026 Alter\nMeridian Pty Ltd (ABN 54 696 662 049).\n\n</details>\n\n---\n\n<div align=\"center\">\n\n<sub><b>~alter</b> is identity infrastructure. Your name is <code>~yourname</code> and claiming one is free.</sub>\n\n<sub>\n<a href=\"https://truealter.com\">Website</a> &nbsp;·&nbsp;\n<a href=\"https://truealter.com/docs\">Docs</a> &nbsp;·&nbsp;\n<a href=\"https://truealter.com/origin\">The argument in eight chapters</a> &nbsp;·&nbsp;\n<a href=\"https://datatracker.ietf.org/doc/search/?name=draft-morrison&activedrafts=on\">The open specifications</a> &nbsp;·&nbsp;\n<a href=\"https://github.com/true-alter\">Every repository</a>\n</sub>\n\n</div>\n",
  "bytes": 20949,
  "sha": "8cbf7e37c0bca40c595cdd382f0a98d0de02e0ab6cbc89b441fd0ce88606bc84",
  "repo_slug": "true-alter/mcp-ollama",
  "fonte": "repo",
  "truncated": false,
  "api": "https://agentalog.com/api/listings/mcp_io_github_true_alter_mcp_ollama_1e28c581/readme"
}