io.github.batchwatch/batchwatch-mcp
Ask batchwatch whether to batch a job and get the real measured queue evidence, not a guess.
Open source Open in the app JSON README (API)
About
Ask batchwatch whether to batch a job and get the real measured queue evidence, not a guess.
Details
- Kind
- MCP servers
- Topic
- No topic detected
- Publisher
- batchwatch
- Origin
- official
- Category
- ferramentas
- Transport
- local
- Version
- 0.1.1
- Last push
- 2026-08-31T11:44:23Z
- Repository state
- ativo
- Language
- C++
- Added
- 2026-08-31 12:00:08
- Updated
- 2026-08-31 12:00:08
- Origin id
io.github.batchwatch/batchwatch-mcp
README
# Client libraries
batchwatch only works if people instrument their pipelines, and nobody writes
raw HTTP calls to donate data. These are the libraries that make it two
lines.
| Package | Language | Tests | State |
|---|---|---|---|
| [`python/`](python) | Python 3.8+ | 96, all passing | works; PyPI release on the way |
| [`typescript/`](typescript) | TypeScript / JS, Node 20+ | 84, all passing (built + tested in CI) | works; built and tested in CI; npm release on the way |
| [`dotnet/`](dotnet) | C#, `net8.0` | built and tested in CI (xunit) | works; NuGet release on the way |
| [`go/`](go) | Go 1.21+ | 83, all passing (incl. `-race`) | works; module-proxy release on the way |
| [`ruby/`](ruby) | Ruby 3.0+ | 87, all passing (minitest) | works; RubyGems release on the way |
| [`php/`](php) | PHP 8.2+ | all passing | works; on Packagist, tagged release on the way |
| [`java/`](java) | Java 17+ | 88, all passing | works; Maven Central release on the way |
| [`rust/`](rust) | Rust 1.63+ | 94, all passing (`cargo test`) | works, std-only; http-only (see note); crates.io release on the way |
| [`cpp/`](cpp) | C++17 (POSIX) | 85, all passing | works, stdlib+sockets only; source-only by design; http-only (see note) |
All nine expose the **same surface**, and a [conformance check](conformance) fails CI if
any language falls behind — see [The same surface, in all nine](#the-same-surface-in-all-nine).
`client/batchwatch.py` in the repo root is the original single-file client
and is left untouched. `clients/python/` is the packaged version of it, plus
spooling.
**A note on TLS (Rust and C++).** Every client except Rust and C++ gets
TLS from its standard library and talks to `https://batchwatch.dev` directly.
Rust's and C++'s standard libraries have no TLS, and both are written with
zero external dependencies on purpose, so their transport is **`http://`
only**. Pointed at the default `https://` URL they *spool* rather than
deliver (the measurement is kept, not lost) until you point them at an
`http://` endpoint or a local TLS-terminating proxy. Their READMEs say so.
## What every client does the same way
**It fails open.** A batchwatch outage must never stop a user's job. Every
submission happens off the caller's thread with a short timeout, every error
is swallowed and logged at debug level, and the only call you await —
`should_batch()` — returns *your* default when it cannot answer, never a
guess. The default is "run it synchronously": being wrong that way costs
money, being wrong the other way blows a deadline. Each package has a test
that runs against a dead port and a hung socket.
**Two lines to adopt.** `should_batch()` before you submit, `track()` around
the call.
**It never sends content.** No prompts, no completions, no file names. The
body is built from one allowlist — provider, model, mode, endpoint, request
count, token counts, timestamps, status — and everything else is dropped by a
single function - `_scrub` in Python, `Scrub` in Go, `clean` in TypeScript,
`sanitize` in PHP, Ruby and Java, `strip` in Rust and C++ - on the way out.
Each package has
a test that asserts this on what the server actually received, with a
positive control so it cannot pass by sending nothing at all.
**`output_tokens` defaults to null, never 0.** Output costs five to six times
as much as input, so a saving computed on zero output is systematically too
low — 3.4x too low in the case that led to this rule — and nothing in the
response reveals it. Absence must stay absence all the way to the server.
Explicitly passing `0` still sends `0`: zero is a measurement.
**It spools to disk.** An undeliverable *completed* measurement is appended
to a JSONL file and replayed later via `POST /v1/calls/complete`. Losing
measurements when the network is bad means losing them exactly when they are
most interesting.
**It does the annoying parts.** Beyond the two-line advisory path, every client
carries the same high-level surface so you never hand-roll it:
- **The high-level batch job** — `batch(...)` hands the client the two callables
(batch-create + a synchronous fallback) and it owns the rest: a **deadline
guard** that shifts to the fallback when the wait runs long, a **poll loop**
with exponential backoff, jitter, a rate-limit floor and a first cadence
informed by the model's measured p50, and **partial-completion** handling that
splits a batch into landed / failed / expired mapped by `custom_id` (never by
index) with an idempotent retry of only the failed subset. We take the
callable, never the payload — the deadline fallback is reported down the same
accuracy path a completion uses, so nothing new is sent.
- **Read your own contributions** — `my_calls()` and `key_status()`
(`GET /v1/calls/mine`, `/v1/keys/current`): the per-key readback for verifying
a measurement landed and checking your tier/quota.
- **Subscribe to outage alerts** — `subscribe()` / `subscriptions()` /
`unsubscribe()` against `/v1/subscriptions`, the "own the outage moment"
channel.
Unlike the measurement path, these last two do **not** fail open: they are
explicit actions against a per-key route, so without a key they raise rather than
silently pretend. The job path is the user's own job, so a misuse (a result
before a submit, a deadline with no fallback) raises loudly too — only telemetry
fails open.
**No dependencies.** Standard library only, in all nine.
## The same surface, in all nine
Every client promises the same capabilities, and — since a feature can land in
one language, its card be closed in good faith, and the other eight silently lag
— [`conformance/`](conformance) is the check that stops that. `manifest.json`
declares the promised surface; `check.py` greps each SDK's own source **and its
own tests** (per language, never a loose cross-language match) and fails CI on
any unexplained gap. Exemptions must be explicit and justified in the manifest.
Run it with `python clients/conformance/check.py --list`.
## The spool format
One JSON object per line, in the shape `/v1/calls/complete` accepts:
```json
{"provider":"openai","model":"gpt-5.6-sol","mode":"batch","requests":1,
"endpoint":null,"input_tokens":9720,"output_tokens":null,"status":"completed",
"started_at":"2026-08-25T10:00:00Z","ended_at":"2026-08-25T10:04:00Z"}
```
Identical across all nine clients, so a file written by one can be flushed
by another. Default location is `$BATCHWATCH_SPOOL`, otherwise
`batchwatch-spool.jsonl` in the temp directory.
Two consequences worth knowing before you rely on it:
- **Spooling needs an API key.** `/v1/calls/complete` takes the caller's own
timestamps, so it is closed to anonymous callers — see the reasoning in
`src/index.js`. A client without a token therefore does not spool at all:
a file that can never be sent is a disk leak, not data safety.
- **Replay can duplicate.** If the original `PATCH` reached the server but
the response did not, the spooled copy arrives as a second row. That is the
deliberate trade: a duplicate is visible in the dataset, a lost measurement
is not.
The file is capped (5 MB by default). Past the cap, measurements are dropped
rather than filling the user's disk.
## Publishing
Every client works today: install it from the repo (each README shows how) and
it runs. Registry publishing is the next step, and it is in flight — the PHP
package is already on Packagist, the .NET and TypeScript clients build and test
in CI on every push, and the rest are being wired up (PyPI, npm, a Go module
proxy, RubyGems, Maven Central, crates.io). Client CI itself lands in
[#184](https://github.com/batchwatch/client/issues/184).
Two properties are deliberate design decisions, not gaps, and each has a
workaround in the relevant README:
- **Rust and C++ speak `http://` only** — their standard libraries carry no TLS
and both are zero-dependency by design. Point them at a local TLS-terminating
proxy in front of `batchwatch.dev`, and they deliver directly (see the TLS
note above).
- **The POSIX-socket clients (C++) are first-class on Linux and other POSIX
platforms;** Windows needs a Winsock shim.