{
  "markdown": "<p align=\"left\"><img src=\"https://raw.githubusercontent.com/ShekharBhardwaj/AgenticLedger/main/docs/raccoon.svg\" alt=\"\" width=\"72\" height=\"66\"></p>\n\n# Agentic Ledger\n\n[![CI](https://github.com/ShekharBhardwaj/AgenticLedger/actions/workflows/ci.yml/badge.svg)](https://github.com/ShekharBhardwaj/AgenticLedger/actions/workflows/ci.yml)\n[![CodeQL](https://github.com/ShekharBhardwaj/AgenticLedger/actions/workflows/codeql.yml/badge.svg)](https://github.com/ShekharBhardwaj/AgenticLedger/actions/workflows/codeql.yml)\n[![PyPI](https://img.shields.io/pypi/v/agentic-ledger)](https://pypi.org/project/agentic-ledger/)\n[![Python versions](https://img.shields.io/pypi/pyversions/agentic-ledger)](https://pypi.org/project/agentic-ledger/)\n[![Docker](https://img.shields.io/badge/docker-ghcr.io-blue)](https://ghcr.io/shekharbhardwaj/agentic-ledger)\n[![License: MIT](https://img.shields.io/badge/license-MIT-green)](LICENSE)\n[![MCP server on Glama](https://glama.ai/mcp/servers/ShekharBhardwaj/AgenticLedger/badges/score.svg)](https://glama.ai/mcp/servers/ShekharBhardwaj/AgenticLedger)\n[![BYOAIK status](https://qhfc1deef2.execute-api.us-east-1.amazonaws.com/tools/agentic-ledger/badge.svg)](https://www.byoaik.com/tools/agentic-ledger/)\n\nRuntime observability for AI agents — see exactly what your agent did, why it did it, and what it cost.\n\n**Website:** [agentic-ledger.dev](https://agentic-ledger.dev)\n\n> The numbers are meant to match your provider bill. If they don't, [that's a bug we want](https://github.com/ShekharBhardwaj/AgenticLedger/issues/new/choose).\n\nWorks with **any agent framework**, **any LLM provider**, **any model gateway**. Zero code changes required. Point your agent at the proxy and everything is captured automatically.\n\n---\n\n## How it works\n\nAgentic Ledger runs as a transparent proxy between your agent and the LLM provider. It intercepts every request and response, assigns it an `action_id`, stores it, and returns the upstream response unmodified. Your agent never knows the proxy is there. The full picture, with\ndiagrams and a module map for contributors, lives in\n[ARCHITECTURE.md](ARCHITECTURE.md).\n\n```\nYour Agent  →  Agentic Ledger Proxy  →  OpenAI / Anthropic / LiteLLM / any LLM\n                      ↓\n               SQLite or Postgres\n                      ↓\n               Live Dashboard + API\n```\n\n---\n\n## Quick Start\n\n**Step 1 — Start the proxy**\n\nTwo commands, zero config, no terminal held hostage:\n\n```bash\nuv tool install agentic-ledger    # or: pipx install agentic-ledger, or pip install -U agentic-ledger\nagenticledger start     # runs in the background; terminal freed\n```\n\nA tool-managed install (`uv tool` / `pipx`) gets its own isolated\nenvironment and one unambiguous shim on PATH, so it can never be\nshadowed by another Python's copy and `agenticledger upgrade` always\nmeans exactly one thing. Plain pip works too; if a machine ever grows\ncompeting installs, `agenticledger doctor --fix` untangles them.\n\n`agenticledger start` prints the dashboard URL and gives your terminal\nback — closing the window doesn't stop it. `agenticledger status` tells\nyou it's up and healthy, `agenticledger logs` shows what it's doing,\n`agenticledger stop` shuts it down. Want a config file anyway?\n`agenticledger init` writes a commented one; see\n[Configuration](#configuration) for what goes in it.\n\nOr with Docker (no Python required):\n```bash\ndocker run -p 8000:8000 \\\n  -e AGENTICLEDGER_UPSTREAM_URL=https://api.openai.com \\  # optional: omit to route by call format\n  -v $(pwd)/data:/data \\\n  ghcr.io/shekharbhardwaj/agentic-ledger:latest\n```\n\n> The image is multi-arch (amd64/arm64), runs as a non-root user, and every\n> release is signed with Sigstore and ships an SBOM. Hardening a shared\n> deployment (TLS, auth keys, redaction, verification)? See the\n> [deployment guide](docs/deployment.md).\n\n> **Using Anthropic / Claude?** Nothing to configure: with no upstream\n> set, the proxy routes each call by its wire format, so Anthropic-style\n> calls go to Anthropic and OpenAI-style calls go to OpenAI, side by side\n> through one proxy. Setting an explicit `upstream_url` (a gateway like\n> LiteLLM or OpenRouter, LM Studio, or a pinned provider) switches to the\n> classic one-proxy-one-provider behavior, mismatch hints included.\n\nOr with docker compose (SQLite by default — see `docker-compose.yml`):\n```bash\nAGENTICLEDGER_UPSTREAM_URL=https://api.openai.com docker compose up\n```\n\nWith `uv`:\n```bash\nuv add agentic-ledger\nAGENTICLEDGER_UPSTREAM_URL=https://api.openai.com uv run python -m agenticledger.proxy\n```\n\nWith `pip`:\n```bash\npython -m venv venv && source venv/bin/activate\npip install -U agentic-ledger\nAGENTICLEDGER_UPSTREAM_URL=https://api.openai.com ./venv/bin/python -m agenticledger.proxy\n```\n\n> **Postgres?** Install the extra and set `AGENTICLEDGER_DSN`:\n> ```bash\n> pip install \"agentic-ledger[postgres]\"\n> AGENTICLEDGER_DSN=postgresql://user:password@localhost/agenticledger\n> ```\n> Note: the Docker image uses SQLite only. For Postgres with Docker, install via `pip` instead.\n\n> **OpenTelemetry?** Install the extra and set `AGENTICLEDGER_OTEL_ENDPOINT`:\n> ```bash\n> pip install \"agentic-ledger[otel]\"\n> AGENTICLEDGER_OTEL_ENDPOINT=http://localhost:4318\n> ```\n\nProxy starts on `http://localhost:8000`. Traces are saved to `~/.agenticledger/agenticledger.db` when started with `agenticledger start` (one home for the background service, wherever you launched it from), to `agenticledger.db` in the current folder when run in the foreground (`agenticledger serve` / `python -m agenticledger.proxy`), or to `/data/agenticledger.db` in Docker.\n\n---\n\n**Step 2 — Point your agent at the proxy**\n\nFor Claude Code, BMAD, or OpenClaw, one command writes the config for you\n(backed up, merged, Docker-aware):\n\n```bash\nagenticledger connect claude-code    # or: bmad, openclaw\n```\n\nFor everything else, two changes: set `base_url` to the proxy and add a session ID header to group calls into a run. Everything else — your API key, model, messages — stays exactly the same.\n\n**OpenAI:**\n```python\nfrom openai import OpenAI\n\nclient = OpenAI(\n    base_url=\"http://localhost:8000/v1\",  # ← proxy\n    api_key=\"your-openai-key\",\n    default_headers={\"x-agenticledger-session-id\": \"run-1\"},\n)\n\nresponse = client.chat.completions.create(\n    model=\"gpt-4o\",\n    messages=[{\"role\": \"user\", \"content\": \"Research the top 3 AI trends in 2026\"}],\n)\n```\n\n**Anthropic** (no upstream config needed: `/v1/messages` calls route to Anthropic automatically):\n```python\nimport anthropic\n\nclient = anthropic.Anthropic(\n    base_url=\"http://localhost:8000\",  # ← proxy\n    api_key=\"your-anthropic-key\",\n    default_headers={\"x-agenticledger-session-id\": \"run-1\"},\n)\n```\n\n**Azure OpenAI:** point `AzureOpenAI(azure_endpoint=\"http://localhost:8000\")` at the ledger with your resource set as the upstream; deployments are priced from the model the response names. See the [Azure guide](docs/integrations/azure-openai.md).\n\n**AWS Bedrock:** install `agentic-ledger[bedrock]`, give the ledger AWS credentials through the standard chain, and point `boto3` (`endpoint_url`) or Claude Code (`ANTHROPIC_BEDROCK_BASE_URL`) at it; the ledger re-signs each call itself. Both wires are covered: InvokeModel and the modern Converse/ConverseStream APIs. See the [Bedrock guide](docs/integrations/bedrock.md).\n\n**LiteLLM / OpenRouter / any gateway:**\n```bash\n# Point Agentic Ledger at your gateway\nAGENTICLEDGER_UPSTREAM_URL=http://localhost:4000 uv run python -m agenticledger.proxy\n\n# Then point your agent at Agentic Ledger\nclient = OpenAI(base_url=\"http://localhost:8000/v1\", ...)\n```\n\n---\n\n**Step 3 — Open the dashboard**\n\n```\nhttp://localhost:8000\n```\n\nThe web app updates live via WebSocket as calls come in. No refresh needed.\n\n- **Loop Lens** — every loop run with status (`running` / `flagged` / `complete` / `ended` / `stopped`), a cost-per-iteration chart, a **block-calls button** that refuses a running loop's further calls at the wall (and an allow-calls-again to lift it; the agent being blocked cannot), per-iteration breakdowns, and plain-English explanations of every flag. Pick any two runs with **⇆** to diff them side by side — cost, iterations, calls, flags, duration with signed deltas, plus a **prompt drift** diff showing exactly what changed in the system prompt and opening instruction between the runs.\n- **Sessions** — every session with three views: expandable call cards (response, thinking, tool calls, cache tokens, interaction badges), a **Flow** DAG of agent handoffs, and a **Trace** waterfall with real parent links from the loop engine. Session cards say whose they are and how they're doing at a glance: team badge, red \"N failed\" for real errors, amber \"N blocked\" for budget walls, purple tiles for replays. Hover a card for one-click delete.\n- **Replay the whole run** — the question that decides a model switch isn't \"how did it handle one call?\" but \"would my loop have survived?\" Pick a run or session, pick a destination (a local model is free), and every step re-runs with its original inputs. You get a report card, not homework: **\"34 / 40 moments matched\"**, the fumbles named (\"dropped the tools\"), and the cost both ways. Each step is a real captured moment replayed honestly — after step one a different model would have steered a different conversation, so the ledger compares moments, not fairy tales.\n- **In your pocket** — `agenticledger share` opens an https tunnel you own (via cloudflared, no account), prints the pairing link, and draws a QR in the terminal: point your phone's camera and the dashboard is in your hand, kill switch and ceilings included. `--wifi` for a same-network link, `--rotate` to un-pair every device, or press **Pair a device** in the dashboard's ⚿ panel. Local machines never need a key; everyone else meets the auto-generated pairing key. The dashboard fits a phone: one pane at a time, a back button, prev/next arrows to flip between runs.\n- **Named instances** — `agenticledger start --name demo --port 8003` runs a second ledger beside your everyday one: own state, own database, its dashboard wears an amber name chip so it can never pass for the real thing. `stop`, `status`, `logs`, `share`, and `run` all take `--name`.\n- **The spend meter** — a run's detail reads its money live: spent so far, burning $/h, \"at this pace $Y by 8:00 AM\". Give any run a **cost ceiling** and the proxy refuses further calls the moment spend reaches it (amber, costing nothing) until you raise or clear it; the ceiling survives restarts and guards auto-detected loops too. A webhook alert fires at 80%.\n- **Names, pins, projects** — call it \"the overnight auth fix\" instead of `cc-73a26366`, ★ pin what matters to the top, file work under a project and the Sessions view reads as sections: a heading per project, its sessions beneath, the unfiled pile last. A run filed under a project files its sessions with it.\n- **Settings** — the ⚙ shows what the proxy is actually running with: config file in effect, upstream, budgets, replay targets, each row labeled file / env / default. Read-only, secrets hidden.\n- **Replay & what-if** — open any call and **↻ Replay** it: pick a destination (the panel lists what your local server actually has loaded), and the exact captured prompt re-executes there — same provider, the other one, or **a free local model via LM Studio**; tool calls, schemas, and system prompts are translated between the Anthropic and OpenAI wire formats automatically. Works even on calls your own budget blocked — the wall can say no and you can still see what would have happened, for $0. Replays tie back to their original with **↩ Open original**. The **what-if** box answers the cheaper question first: reprice any run or session on another model with pure math, no API calls. (Configure `AGENTICLEDGER_REPLAY_API_KEY` and/or the per-provider `AGENTICLEDGER_REPLAY_*_KEY` targets.)\n- **Reports** — where the money goes: spend per day, model mix with **latency p50/p95/p99**, per-agent totals, a **by-team table** with each team's spend against its card's daily allowance (\"who ran dry?\" in one glance), and **cache savings** — what your prompt-cache traffic would have cost at full input rates versus what it actually cost. Errors and blocks are counted apart everywhere: **red = something broke, amber = the ledger refused on purpose** — a healthy wall never makes a healthy agent look sick\n- **Search** — full-text search across all sessions by prompt, output, agent name, or user ID\n\n---\n\n## Configuration\n\n`agenticledger init` writes `agenticledger.toml` with every option\ncommented. Uncomment what you need — a working setup looks like this:\n\n```toml\n[proxy]\nport = 8000\nupstream_url = \"https://api.anthropic.com\"\ndb = \"sqlite:///agenticledger.db\"\n\n[keys]\n# Prefer *_file: the file's contents are the key, so no secret lives in\n# this file or your shell history (chmod 600 the key file).\napi_key_file = \"~/.agenticledger/api.key\"       # dashboard/admin access\ningest_key_file = \"~/.agenticledger/ingest.key\" # closes the open relay\n\n[budgets]\ndaily = 25.0          # whole-ledger daily ceiling, USD\nsession = 5.0         # per-session ceiling\n\n[replay]\n# Free local replay via LM Studio (any key works there):\nopenai_url = \"http://localhost:1234\"\nopenai_key = \"lm-studio\"\n```\n\nThree rules:\n\n1. **The file is found in this order:** `AGENTICLEDGER_CONFIG`, then\n   `./agenticledger.toml` (the folder you start from), then\n   `~/.agenticledger/config.toml`. First match wins; the startup banner\n   names the file in effect.\n2. **Anything typed in the command beats the file.** Env vars override\n   per-setting (`AGENTICLEDGER_PORT=9000 agenticledger start` uses 9000\n   for that run without touching the file) — which is also why Docker and\n   CI setups configured by env vars are unaffected.\n3. **Changes apply on restart** (`agenticledger stop` then `start`).\n\nEvery setting in the [environment-variable reference](#configuration-reference)\nbelow has a config-file home; an `[env]` section passes any other\n`AGENTICLEDGER_*` variable through verbatim.\n\n---\n\n## Providers, step by step\n\nEvery provider below rides the same proxy; the only thing that changes is\nwhich base URL you point at it. Each recipe assumes the proxy is up\n(`agenticledger start`) and ends with the same check: run one call, open\nhttp://localhost:8000, and see it in Sessions.\n\n**OpenAI (and any OpenAI-compatible API)**\n\n1. Point the client at the proxy:\n   ```bash\n   export OPENAI_BASE_URL=http://localhost:8000/v1\n   ```\n2. Keep your `OPENAI_API_KEY` exactly as it was — the proxy passes your\n   auth header through untouched.\n3. Make a call; it appears in Sessions with an O mark.\n\n**Anthropic**\n\n1. Point the client at the proxy:\n   ```bash\n   export ANTHROPIC_BASE_URL=http://localhost:8000\n   ```\n2. Keep your `ANTHROPIC_API_KEY` as it was.\n3. Make a call; it appears with an A mark. No upstream config needed —\n   the proxy routes Anthropic-shaped calls to Anthropic by wire format.\n\n**AWS Bedrock (direct capture)**\n\n1. Give the ledger AWS credentials of its own through the standard chain\n   (env vars, `~/.aws` profile, or an instance role) scoped to\n   `bedrock:InvokeModel` and `bedrock:InvokeModelWithResponseStream`,\n   then install the extra and restart:\n   ```bash\n   pip install \"agentic-ledger[bedrock]\"\n   agenticledger stop && agenticledger start\n   ```\n2. Check the ⚙ Settings panel: the Bedrock row should read \"signing as\n   the ledger in <your region>\".\n3. Point the client at the proxy — Claude Code:\n   ```bash\n   export CLAUDE_CODE_USE_BEDROCK=1\n   export ANTHROPIC_BEDROCK_BASE_URL=http://localhost:8000\n   ```\n   boto3: `boto3.client(\"bedrock-runtime\", endpoint_url=\"http://localhost:8000\")`.\n4. Make a call; it appears with an orange B mark. The ledger strips the\n   caller's identity and re-signs with its own credentials. Full guide:\n   [docs/integrations/bedrock.md](docs/integrations/bedrock.md).\n\n**Azure OpenAI**\n\n1. Set the upstream to your resource:\n   ```bash\n   agenticledger config set proxy.upstream_url https://<resource>.openai.azure.com\n   agenticledger stop && agenticledger start\n   ```\n2. Point the client's Azure endpoint at `http://localhost:8000`; keep\n   your `api-key` header as it was.\n3. Calls are tagged `azure-openai` and priced by the model the RESPONSE\n   names, so deployment aliases can't hide the real model. Full guide:\n   [docs/integrations/azure-openai.md](docs/integrations/azure-openai.md).\n\n**Local models (LM Studio, Ollama with the OpenAI API)**\n\n1. Set the upstream to the local server:\n   ```bash\n   agenticledger config set proxy.upstream_url http://localhost:1234\n   agenticledger stop && agenticledger start\n   ```\n2. `export OPENAI_BASE_URL=http://localhost:8000/v1` in the agent.\n3. Calls appear with a purple mark and $0 cost. Full guide:\n   [docs/integrations/lm-studio.md](docs/integrations/lm-studio.md).\n\n**Gateways (OpenRouter, LiteLLM)**\n\n1. Set the upstream to the gateway:\n   ```bash\n   agenticledger config set proxy.upstream_url https://openrouter.ai/api\n   agenticledger stop && agenticledger start\n   ```\n2. `export OPENAI_BASE_URL=http://localhost:8000/v1`; keep the gateway\n   key as it was.\n3. Gateway-prefixed model ids (\"anthropic/claude-...\") price correctly\n   via substring matching. Guides: [openrouter.md](docs/integrations/openrouter.md),\n   [litellm.md](docs/integrations/litellm.md).\n\nFramework-specific recipes (CrewAI, LangGraph, AutoGen, Vercel AI SDK,\npydantic-ai, and more) live in [docs/integrations/](docs/integrations/).\n\n---\n\n## Coding agents — Claude Code, Ralph loops & friends\n\nClaude Code (and most coding agents) can be pointed at the proxy with a single\nenvironment variable — no headers, no code changes:\n\n```bash\nagenticledger start\n```\n\n```bash\nexport ANTHROPIC_BASE_URL=http://localhost:8000\nclaude\n```\n\nNo upstream config needed: calls route to the provider matching their\nwire format.\n\nAgentic Ledger fingerprints Claude Code traffic automatically: every call is\ntagged `framework=claude-code`, and instead of one undifferentiated bucket,\neach Claude Code session appears under its **real session UUID** (the same id\n`claude --resume` shows), with prompt-cache reads/writes captured and priced\ncorrectly — cache traffic is where most of a coding agent's real spend lives.\n\nWant a loop filed under a name you chose? Put one word in front of the\ncommand you already run:\n\n```bash\nagenticledger run nightly-digest -- python agent.py\n```\n\nYour command runs exactly as before; its LLM calls land on the run tile\nnamed `nightly-digest`, and each launch counts as the next iteration, so\ntomorrow's run joins the same tile. Nothing in your agent's code changes.\nAdd `--project acme` to file the run under a dashboard project as it starts.\n\nRunning an overnight loop (Ralph-style `while :; do cat PROMPT.md | claude -p; done`)?\nThe same command with loop flags re-executes your command each iteration,\nattributes every call to the run (via the base URL, no headers needed), and\nstops on a completion promise, a budget ceiling, or the iteration cap:\n\n```bash\nAGENTICLEDGER_UPSTREAM_URL=https://api.anthropic.com \\\nAGENTICLEDGER_COMPLETION_PROMISE=\"ALL TASKS COMPLETE\" \\\nuv run python -m agenticledger.proxy\n```\n\n```bash\nagenticledger run overnight --max-iterations 50 --budget 25 -- \\\n  claude -p \"$(cat PROMPT.md)\" --dangerously-skip-permissions\n```\n\nEach iteration shows up as *iteration N of the run* in `/api/runs`; when the\nagent prints the completion promise in a response, run status flips to\n`complete` and the loop exits with a cost/token summary. The word after\n`run` is the run's name; without one the run is named after the folder and\nthe minute (`myproject-0819-1936`). Rerunning the same name continues its\niteration count instead of restarting at 1. Any existing loop\nscript works too — poll `GET /api/runs/{run_id}` yourself, or let the proxy's\nbudgets (`AGENTICLEDGER_BUDGET_DAILY=25.00`) hard-stop a runaway loop.\n\nIterating on the prompt? Rerun and use **⇆ compare** in the Loop Lens to diff\nthe two runs — cost, iterations, calls, and flags side by side — so \"did the\nnew prompt actually help\" gets a number instead of a feeling.\n\nThe same recipe works for any client with a base-URL override (Codex CLI,\nopencode, OpenClaw, LiteLLM-based stacks) — set the OpenAI/Anthropic base URL\nto the proxy and traffic is captured; add `x-agenticledger-*` headers when you\nwant explicit attribution.\n\n**OTel-native tools** (Gemini CLI, Codex `[otel]`, AutoGen/AG2, Pydantic AI,\nVercel AI SDK) don't need the proxy at all — point their OTLP exporter at the\nledger and GenAI spans are ingested directly:\n\n```bash\nexport OTEL_EXPORTER_OTLP_ENDPOINT=http://localhost:8000\n```\n\nBoth OTLP/HTTP encodings are accepted: JSON always, protobuf when the\n`[otel]` extra is installed (the Docker image includes it). gRPC exporters\nshould switch to HTTP: `OTEL_EXPORTER_OTLP_PROTOCOL=http/protobuf`.\n\n**Framework guides** — one per integration in\n[docs/integrations](docs/integrations/README.md): Claude Code, Codex CLI,\nopencode, OpenClaw, BMAD-METHOD, LangGraph/LangChain, CrewAI, OpenAI Agents\nSDK, Gemini CLI, AutoGen/AG2, Pydantic AI, Vercel AI SDK, LiteLLM,\nOpenRouter, and LM Studio (fully offline: local model, local ledger).\n\n**Production deployment** — TLS termination, auth keys, redaction, image\nsignature/SBOM verification, enterprise mirrors, and scaling guidance in\n[docs/deployment.md](docs/deployment.md).\n\n---\n\n## The numbers\n\nMeasured, not promised. Reproduce them with\n`python scripts/loadtest.py --calls 2000 --seed 1000000` (Apple M-series\nMacBook, SQLite backend; re-measured on 0.10 with the provider adapter\narchitecture in place — same numbers, 2,580 to 2,800 calls/s on both):\n\n| What | Result |\n|---|---|\n| Sustained capture throughput | 2,886 proxied calls/sec |\n| Added latency per call | 10ms p50 · 12ms p95 |\n| Direct store writes | ~38,000 saves/sec |\n| One million calls on disk | 271 MB |\n| Open one session at 1M calls | 2 ms |\n| Session list at 1M calls | 335 ms |\n| 30-day report at 1M calls | 719 ms |\n\nThe honest caveats: the session list aggregates every session on every\nload, so it grows with total history; the report window uses a timestamp\nindex, so it grows with the window's traffic, not the table. Your agent's\nprovider latency (hundreds of ms per call) dwarfs the proxy's overhead by\nan order of magnitude. Postgres numbers vary with your server; the same\nscript measures them with `--dsn`. Cost math has its own guardrails and\na five-minute parity check against your provider console: see\n[docs/accuracy.md](docs/accuracy.md).\n\n## What gets captured\n\nEvery LLM call is stored with:\n\n| Field | What it contains |\n|---|---|\n| `action_id` | UUID assigned at interception time |\n| `session_id` | Run grouping (from header) |\n| `timestamp` | When the call was made |\n| `model_id` | Model used |\n| `provider` | `openai` or `anthropic` |\n| `messages` | Full message history sent to the model |\n| `system_prompt` | Extracted system prompt |\n| `tools` | Tool definitions available to the model |\n| `tool_calls` | Tools the model decided to call |\n| `tool_results` | What the tools returned (from next call's messages) |\n| `content` | Model's text output |\n| `stop_reason` | Why the model stopped |\n| `tokens_in` / `tokens_out` | Token usage |\n| `cache_read_tokens` / `cache_write_tokens` | Prompt-cache usage — reads and writes are priced correctly per provider |\n| `thinking` | Extended-thinking output (Anthropic), captured separately from `content` |\n| `cost_usd` | Estimated cost based on model pricing |\n| `latency_ms` | End-to-end response time |\n| `status_code` | HTTP status from upstream — errors are captured too |\n| `error_detail` | Upstream error message for non-200 responses |\n| `agent_name` | From `x-agenticledger-agent-name` header, or auto-detected (e.g. `claude-code`) |\n| `framework` | From `x-agenticledger-framework` header, or fingerprint-detected (e.g. `claude-code`, `gemini-cli`, `litellm`) |\n| `user_id` | From `x-agenticledger-user-id` header |\n| `app_id` | From `x-agenticledger-app-id` header |\n| `environment` | From `x-agenticledger-environment` header |\n| `parent_action_id` | Parent call in a nested agent graph |\n| `handoff_from` / `handoff_to` | Agent handoff tracking for the Flow DAG |\n\n---\n\n## API reference\n\n| Method | Endpoint | Description |\n|---|---|---|\n| `GET` | `/health` | Liveness — `{\"status\":\"ok\",\"version\":\"...\"}`. No auth, never touches the store. |\n| `GET` | `/readyz` | Readiness — pings the store; `503` when unreachable. Also reports `capture_dropped`. |\n| `GET` | `/metrics` | Prometheus metrics (captures persisted/dropped, queue depth). |\n| `GET` | `/api/audit` | Audit trail of sensitive actions (admin). |\n| `DELETE` | `/api/users/{user_id}` | Right-to-erasure: delete all of a user's captured calls (admin). |\n| `GET` | `/` | Live dashboard |\n| `WS` | `/ws` | WebSocket stream — powers live dashboard updates |\n| `GET` | `/api/sessions` | List recent sessions with aggregated stats |\n| `GET` | `/api/runs` | List loop runs (explicit or auto-inferred) with iterations, cost, status, and flagged-call counts |\n| `GET` | `/api/runs/{run_id}` | One run's status (`running` / `flagged` / `complete` / `ended` / `stopped`) — poll this from loop scripts |\n| `GET` | `/api/sessions/{session_id}/tools` | Derived tool executions — each tool call paired with its result, latency, and error status |\n| `DELETE` | `/api/sessions/{session_id}` | Delete a session and all its calls |\n| `GET` | `/api/reports?days=30` | Spend insights: daily trend, model mix with signed cache savings, latency percentiles, per-agent and per-team totals |\n| `GET` | `/api/whatif?model=...&run_id=...` | Reprice a run/session/call's captured tokens on another model — pure math, zero API calls |\n| `POST` | `/api/tokens` | Mint scoped API tokens — including `role: ingest` team cards with `budget_daily` |\n| `GET` | `/api/calls/{action_id}` | One call by id — follow a replay's parent back to its original |\n| `GET` | `/api/replay/targets` | Configured replay destinations (feeds the dashboard's dropdown) |\n| `GET` | `/api/replay/models` | Models a replay target actually serves (`?provider=`) |\n| `GET` | `/api/whoami` | What is the key I'm holding? Name, role, and team (for team cards) — the dashboard's ⚿ panel uses this |\n| `POST` | `/api/replay/batch` | Replay a whole run or session on another model — returns a job id |\n| `GET` | `/api/replay/jobs/{job_id}` | Batch progress and the report card |\n| `PUT` | `/api/labels/{scope}/{ref_id}` | Name, pin, or file a session/run under a project |\n| `GET` | `/api/projects` | Project names in use |\n| `GET` | `/api/settings` | What the proxy is running with (admin; secrets masked) |\n| `POST` | `/api/redetect` | Re-run framework detection over unattributed history; returns examined, updated, and per-framework counts |\n| `POST` | `/api/replay` | Re-execute a captured call — same provider or translated to the other one (`model` + optional `provider`); result stored linked to the original |\n| `GET` | `/api/search?q=...` | Full-text search across all captured calls |\n| `GET` | `/session/{session_id}` | All calls in a session, ordered by time |\n| `GET` | `/explain/{action_id}` | Single call by action ID |\n| `GET` | `/export/{session_id}` | JSON compliance export with SHA-256 integrity hash |\n| `GET` | `/export/{session_id}/report` | Printable HTML audit report |\n| `POST` | `/mcp` | MCP tool server — `list_sessions`, `explain`, `get_session`, `search`, `list_runs`, `get_run_status` |\n| `POST` | `/v1/traces` | OTLP/HTTP JSON ingest — GenAI spans from OTel-native tools become ledger calls (`/v1/logs` also ingests Claude Code tool events into tool timings; `/v1/metrics` acked) |\n\n**Examples:**\n```bash\n# All calls in a session\ncurl http://localhost:8000/session/run-1\n\n# Search across all sessions\ncurl \"http://localhost:8000/api/search?q=failed+to+connect\"\n\n# Download JSON audit trail (includes an integrity tag; keyed HMAC when configured)\ncurl http://localhost:8000/export/run-1 -o audit-run-1.json\n\n# Printable HTML report — open in browser, print to PDF\nopen http://localhost:8000/export/run-1/report\n```\n\n---\n\n## MCP server\n\nAgentic Ledger exposes its captured data as an MCP (Model Context Protocol) tool server at `POST /mcp`. Point Claude Desktop, Cursor, or any MCP-compatible client at it to query traces directly from your AI assistant.\n\n**Tools available:**\n\n| Tool | Description |\n|---|---|\n| `list_sessions` | List recent sessions with cost, token, and call count summaries |\n| `explain(action_id)` | Full trace for a single LLM call — prompt, tool calls, output, tokens, cost |\n| `get_session(session_id)` | All calls in a session in chronological order |\n| `search(query)` | Full-text search across all captured calls |\n| `list_runs` | Loop runs with iterations, cost, and status |\n| `get_run_status(run_id)` | One run's status — lets an agent inspect its own loop and decide whether to continue |\n\n**Configure in `claude_desktop_config.json`** (HTTP, against a running proxy):\n```json\n{\n  \"mcpServers\": {\n    \"agenticledger\": {\n      \"url\": \"http://localhost:8000/mcp\"\n    }\n  }\n}\n```\n\n**Or as a stdio subprocess** — for clients that launch servers as commands\n(no running proxy required; reads the same database):\n```json\n{\n  \"mcpServers\": {\n    \"agenticledger\": {\n      \"command\": \"agenticledger\",\n      \"args\": [\"mcp\"],\n      \"env\": { \"AGENTICLEDGER_DSN\": \"sqlite:////absolute/path/to/agenticledger.db\" }\n    }\n  }\n}\n```\n\nIf `AGENTICLEDGER_API_KEY` is set, pass it as a header:\n```json\n{\n  \"mcpServers\": {\n    \"agenticledger\": {\n      \"url\": \"http://localhost:8000/mcp\",\n      \"headers\": { \"x-agenticledger-api-key\": \"your-key\" }\n    }\n  }\n}\n```\n\nOnce connected, you can ask your assistant things like:\n- *\"What did the SearchAgent do in the last session?\"*\n- *\"Show me all calls that mentioned rate limit errors\"*\n- *\"What was the total cost of session run-abc123?\"*\n\n---\n\n## Configuration reference\n\nEvery variable below can also live in `agenticledger.toml` — see\n[Configuration](#configuration) for the file, the search order, and the\nenv-always-wins rule.\n\n### Environment variables\n\n**Core:**\n\n| Variable | Required | Default | Description |\n|---|---|---|---|\n| `AGENTICLEDGER_UPSTREAM_URL` | No | _(unset: route by call format)_ | LLM endpoint to forward requests to. Accepts OpenAI, Anthropic, LiteLLM, OpenRouter, or any OpenAI-compatible URL. Omit it and the proxy routes each call by its wire format: Anthropic-shaped calls to Anthropic, Bedrock paths to Bedrock, everything else to OpenAI. |\n| `AGENTICLEDGER_DSN` | No | `sqlite:///agenticledger.db` (Docker: `sqlite:////data/agenticledger.db`) | Database. SQLite for local dev, Postgres URL for production. |\n| `AGENTICLEDGER_HOST` | No | `0.0.0.0` | Host to bind to. Use `127.0.0.1` to restrict to localhost only. |\n| `AGENTICLEDGER_PORT` | No | `8000` | Port to run on. |\n| `AGENTICLEDGER_API_KEY` | No | _(none)_ | Master admin key. When set, the dashboard, read, and management endpoints require authentication; the key grants the `admin` role and bootstraps API tokens (below). Skip for local dev; set when the proxy is on a server — you choose the value. |\n| `AGENTICLEDGER_INGEST_KEY` | No | _(none)_ | When set, the proxy forwards a request only if it carries a matching `x-agenticledger-ingest-key` header — closing the open relay. Off by default; a loud startup warning fires when unset. |\n| `AGENTICLEDGER_REPLAY_API_KEY` | No | _(none)_ | Key for same-provider replay through the proxy's own upstream — the proxy never stores agent credentials, so re-execution needs its own. |\n| `AGENTICLEDGER_REPLAY_OPENAI_KEY` / `_URL` | No | _(none)_ / provider API | Cross-provider replay target: replay **any** capture on OpenAI-format models. Point `_URL` at LM Studio (`http://localhost:1234`, any key) and replaying your captured Claude calls on a local model is **free**. |\n| `AGENTICLEDGER_REPLAY_ANTHROPIC_KEY` / `_URL` | No | _(none)_ / provider API | Cross-provider replay target for Claude models. |\n| `AGENTICLEDGER_*_KEY_FILE` | No | _(none)_ | Every key above also reads from a file named by its `_FILE` variant — the Docker-secrets pattern; keeps keys out of shell history. |\n| `AGENTICLEDGER_EXPORT_HMAC_KEY` | No | _(none)_ | When set, compliance exports carry a tamper-evident keyed `hmac-sha256` integrity tag instead of a plain `sha256` checksum. |\n| `AGENTICLEDGER_EXTRA_PATHS` | No | _(none)_ | Comma-separated additional request paths to capture, e.g. `v1/responses,v1/custom`. Built-in paths (`v1/chat/completions`, `v1/messages`, `v1/responses`, plus `v1/messages/count_tokens` recorded as a free call) are always captured. |\n| `AGENTICLEDGER_ASYNC_CAPTURE` | No | `off` | Persist captures on a background worker so storage never adds latency to the agent's call. Trade-off: reads become **eventually consistent** (a just-captured call may not be queryable for a brief moment). Recommended for high throughput. |\n| `AGENTICLEDGER_CAPTURE_QUEUE_MAX` | No | `10000` | Max captures buffered in async mode before load is shed (drops are counted in `/metrics`). |\n| `AGENTICLEDGER_CAPTURE_LEVEL` | No | `full` | `full` stores everything; `metadata` stores only metrics/metadata (model, tokens, cost, latency, agent, status) and drops prompts, responses, and tools. |\n| `AGENTICLEDGER_REDACT` | No | _(off)_ | Redact PII/secrets in stored data: `all`, or a comma list of `email,ssn,credit_card,ip,api_key`. Replaces matches with `[REDACTED:<label>]`. Only the stored copy is affected — the agent's response is untouched. |\n| `AGENTICLEDGER_REDACT_PATTERNS` | No | _(none)_ | Extra redaction regexes as JSON: `{\"label\": \"regex\", ...}` or `[\"regex\", ...]`. |\n| `AGENTICLEDGER_RETENTION_DAYS` | No | _(keep forever)_ | Delete captured calls older than N days via a background purge worker. |\n| `AGENTICLEDGER_AUDIT_LOG` | No | `on` | Record an audit trail of who viewed/exported/deleted what plus token/erasure actions. Set `0` to disable. |\n\n**Cost budgets** — block calls that exceed a spend limit (returns HTTP 429):\n\n| Variable | Default | Description |\n|---|---|---|\n| `AGENTICLEDGER_BUDGET_SESSION` | _(none)_ | Max USD per `session_id` across its lifetime. |\n| `AGENTICLEDGER_BUDGET_AGENT` | _(none)_ | Max USD per `agent_name` per calendar day (UTC). |\n| `AGENTICLEDGER_BUDGET_DAILY` | _(none)_ | Max USD total across all calls per calendar day (UTC). |\n| `AGENTICLEDGER_BUDGET_USER` | _(none)_ | Max USD per `user_id` per calendar day (UTC) — follows the user across sessions. |\n| `AGENTICLEDGER_BUDGET_STATUS` | `429` | HTTP status for budget blocks. `429` ships with an honest `Retry-After` (seconds until the UTC-midnight window reset); set `402` if your clients retry 429s aggressively — nothing retries Payment Required. |\n| `AGENTICLEDGER_BUDGET_ACTION` | `block` | What happens when a budget is exceeded: `block` returns HTTP 429 (call never reaches the LLM), `warn` lets the call through and fires a webhook alert, `both` blocks and fires the webhook. |\n\n**Rate limits** — block calls that exceed request frequency (returns HTTP 429, sliding 60-second window):\n\n| Variable | Default | Description |\n|---|---|---|\n| `AGENTICLEDGER_RATE_LIMIT_RPM` | _(none)_ | Max requests per minute globally. |\n| `AGENTICLEDGER_RATE_LIMIT_SESSION_RPM` | _(none)_ | Max requests per minute per `session_id`. |\n| `AGENTICLEDGER_RATE_LIMIT_AGENT_RPM` | _(none)_ | Max requests per minute per `agent_name`. |\n| `AGENTICLEDGER_RATE_LIMIT_USER_RPM` | _(none)_ | Max requests per minute per `user_id`. |\n\n**Loop engine** — every call is stitched into ReAct threads (`thread_id`, `step_index`, `prev_action_id`) and fresh-context loop iterations are grouped into runs, with stuck-loop detection:\n\n| Variable | Default | Description |\n|---|---|---|\n| `AGENTICLEDGER_LOOP_ACTION` | `warn` | `warn` records `loop_flags` and fires a `loop_flag` webhook alert; `block` additionally returns HTTP 429 (`loop_detected`) for a session that tripped a guard; `off` disables inference. |\n| `AGENTICLEDGER_LOOP_REPEAT_THRESHOLD` | `3` | Consecutive identical tool calls (same tool, same arguments) before a thread is flagged stuck. |\n| `AGENTICLEDGER_LOOP_MAX_STEPS` | _(none)_ | Flag (and in block mode, stop) threads that exceed this many ReAct steps. |\n| `AGENTICLEDGER_LOOP_RUN_GAP_SECONDS` | `900` | Max gap between fresh-context spawns (same system prompt) that still count as iterations of one run. |\n| `AGENTICLEDGER_COMPLETION_PROMISE` | _(none)_ | Regex matched against response text. On match the call is flagged `completion_promise` and the run's status becomes `complete` — loop runners poll `GET /api/runs/{run_id}` and stop. |\n\n**Alerts** — POST to your webhook when a threshold is breached (does not block calls — see [Alerts](#alerts)):\n\n| Variable | Default | Description |\n|---|---|---|\n| `AGENTICLEDGER_ALERT_WEBHOOK_URL` | _(none)_ | URL to POST alert payloads to. Required for any alerts to fire. |\n| `AGENTICLEDGER_DIGEST_HOUR` | _(off)_ | UTC hour (0–23) to POST a daily spend digest — last 24h totals, cache savings, top models/agents — to the alert webhook. Slack-incoming-webhook friendly (`text`). |\n| `AGENTICLEDGER_ALERT_COST_PER_CALL` | _(none)_ | Alert when a single call costs more than `$X`. |\n| `AGENTICLEDGER_ALERT_LATENCY_MS` | _(none)_ | Alert when a single call takes longer than `Xms`. |\n| `AGENTICLEDGER_ALERT_ERROR_RATE` | _(none)_ | Alert when session error rate exceeds `X` (e.g. `0.5` = 50%). |\n| `AGENTICLEDGER_ALERT_DAILY_SPEND` | _(none)_ | Alert when daily spend crosses `$X`. Unlike budgets, this does not block calls. |\n\n**OpenTelemetry** — emit spans to any OTLP-compatible collector (requires `pip install \"agentic-ledger[otel]\"` — see [OpenTelemetry export](#opentelemetry-export)):\n\n| Variable | Default | Description |\n|---|---|---|\n| `AGENTICLEDGER_OTEL_ENDPOINT` | _(none)_ | OTLP/HTTP base URL, e.g. `http://localhost:4318`. OTel export is disabled when not set. |\n| `AGENTICLEDGER_OTEL_SERVICE_NAME` | `agenticledger` | Value of `service.name` reported to the collector. |\n| `AGENTICLEDGER_OTEL_HEADERS` | _(none)_ | Comma-separated `key=value` auth headers, e.g. `x-honeycomb-team=abc123`. |\n\n**Pricing overrides** — override or extend the built-in per-token pricing table (merged at startup):\n\n| Variable | Default | Description |\n|---|---|---|\n| `agenticledger pricing update` | | Fetch the current price packs from the repository into `~/.agenticledger/pricing/` (overrides built-ins on next start). Network is touched only when you run it (the same is true of `agenticledger upgrade`); nothing phones home on its own. |\n| `AGENTICLEDGER_PRICING` | _(none)_ | Inline JSON map of model → `[input_per_million, output_per_million]` USD. E.g. `'{\"gpt-4o\": [2.50, 10.00], \"my-model\": [1.00, 2.00]}'`. |\n| `AGENTICLEDGER_PRICING_FILE` | _(none)_ | Path to a JSON file with the same format. Applied after `AGENTICLEDGER_PRICING`. |\n\n---\n\n### Common startup examples\n\n```bash\n# Local dev — OpenAI (default)\nAGENTICLEDGER_UPSTREAM_URL=https://api.openai.com uv run python -m agenticledger.proxy\n\n# Local dev — Anthropic\nAGENTICLEDGER_UPSTREAM_URL=https://api.anthropic.com uv run python -m agenticledger.proxy\n\n# Local dev — LiteLLM gateway (any model)\nAGENTICLEDGER_UPSTREAM_URL=http://localhost:4000 uv run python -m agenticledger.proxy\n\n# Production — Postgres + auth + budgets + rate limits + alerts\nAGENTICLEDGER_UPSTREAM_URL=https://api.openai.com \\\nAGENTICLEDGER_DSN=postgresql://user:password@localhost/agenticledger \\\nAGENTICLEDGER_API_KEY=my-secret \\\nAGENTICLEDGER_BUDGET_DAILY=20.00 \\\nAGENTICLEDGER_BUDGET_SESSION=2.00 \\\nAGENTICLEDGER_RATE_LIMIT_SESSION_RPM=20 \\\nAGENTICLEDGER_RATE_LIMIT_USER_RPM=60 \\\nAGENTICLEDGER_ALERT_WEBHOOK_URL=https://hooks.slack.com/services/xxx/yyy/zzz \\\nAGENTICLEDGER_ALERT_COST_PER_CALL=0.50 \\\nAGENTICLEDGER_ALERT_DAILY_SPEND=15.00 \\\nuv run python -m agenticledger.proxy\n```\n\nWhen `AGENTICLEDGER_API_KEY` is set, pass it to access protected endpoints:\n```bash\n# Header\ncurl -H \"x-agenticledger-api-key: my-secret\" http://localhost:8000/session/run-1\n\n# Query param (browser)\nhttp://localhost:8000?api_key=my-secret\n```\n\n#### Scoped API tokens (RBAC)\n\nThe master key is convenient but coarse. For team access, mint **scoped, revocable tokens** with roles instead of sharing the master secret. Tokens are random secrets shown once at creation; only their SHA-256 hash is stored.\n\nRoles are hierarchical:\n\n| Role | Can |\n|---|---|\n| `viewer` | read captured data — dashboard, API, export, MCP |\n| `editor` | viewer + delete sessions |\n| `admin` | editor + manage API tokens |\n\n```bash\n# Mint a viewer token (admin only — use the master key to bootstrap)\ncurl -X POST http://localhost:8000/api/tokens \\\n  -H \"x-agenticledger-api-key: my-secret\" \\\n  -H \"content-type: application/json\" \\\n  -d '{\"name\": \"grafana-readonly\", \"role\": \"viewer\", \"expires_in_days\": 90}'\n# → {\"token_id\": \"...\", \"token\": \"agl_…\", \"role\": \"viewer\", ...}  (token shown once)\n\n# Use it (Bearer header, x-agenticledger-token, or ?token=)\ncurl -H \"Authorization: Bearer agl_…\" http://localhost:8000/api/sessions\n\n# List and revoke\ncurl -H \"x-agenticledger-api-key: my-secret\" http://localhost:8000/api/tokens\ncurl -X DELETE -H \"x-agenticledger-api-key: my-secret\" http://localhost:8000/api/tokens/<token_id>\n```\n\n> Auth is enforced only when `AGENTICLEDGER_API_KEY` is set; the master key is the admin bootstrap for minting tokens. The live `/ws` feed accepts the same credentials (`?api_key=`, `?token=`, `Authorization: Bearer`, or `x-agenticledger-token`) and rejects unauthenticated connects with close code 1008 — the dashboard forwards its page credential to the socket automatically.\n\n---\n\n### Request headers\n\nPass these from your agent on each LLM call. All optional. They enrich captured data, power the Flow tab, and enable per-dimension budgets and rate limits.\n\n| Header | Default | Description |\n|---|---|---|\n| `x-agenticledger-session-id` | _(none)_ | Groups all calls in a run. Use a consistent ID per agent execution (e.g. a UUID or `\"run-1\"`). Without this, calls are stored but not grouped in the dashboard. |\n| `x-agenticledger-user-id` | _(none)_ | End user who triggered this run. Enables per-user rate limiting and auditing. |\n| `x-agenticledger-agent-name` | _(none)_ | Name of the agent making this call (e.g. `\"orchestrator\"`, `\"researcher\"`). Powers the Flow tab DAG and agent-level budgets and rate limits. |\n| `x-agenticledger-app-id` | _(none)_ | Application name or ID. Useful when multiple apps share one proxy. |\n| `x-agenticledger-parent-action-id` | _(none)_ | The `action_id` of the call that spawned this one. When set, the Trace tab draws explicit parent→child connectors. Without it, the Trace tab infers relationships from timestamps automatically. |\n| `x-agenticledger-environment` | `development` | `production`, `staging`, or `development`. Shown in the dashboard. |\n| `x-agenticledger-handoff-from` | _(none)_ | Agent handing off control (e.g. `\"orchestrator\"`). Renders as a directed edge in the Flow DAG. |\n| `x-agenticledger-handoff-to` | _(none)_ | Agent receiving control (e.g. `\"researcher\"`). Renders as a directed edge in the Flow DAG. |\n| `x-agenticledger-framework` | _(auto-detected)_ | Framework/tool making the call (e.g. `\"langgraph\"`, `\"bmad\"`). When absent, well-known clients are fingerprinted automatically (Claude Code, Gemini CLI, LiteLLM). |\n| `x-agenticledger-run-id` | _(auto-inferred)_ | Groups sessions into a loop run (e.g. a Ralph overnight run). When absent, fresh-context sessions sharing a system prompt within `AGENTICLEDGER_LOOP_RUN_GAP_SECONDS` are grouped automatically. |\n| `x-agenticledger-iteration` | _(auto-inferred)_ | Iteration number within the run. |\n\n**Single agent — fully annotated:**\n```python\nfrom openai import OpenAI\n\nclient = OpenAI(\n    base_url=\"http://localhost:8000/v1\",\n    api_key=\"your-openai-key\",\n    default_headers={\n        \"x-agenticledger-session-id\":  \"run-abc123\",\n        \"x-agenticledger-user-id\":     \"user-42\",\n        \"x-agenticledger-agent-name\":  \"researcher\",\n        \"x-agenticledger-app-id\":      \"my-app\",\n        \"x-agenticledger-environment\": \"production\",\n    },\n)\n```\n\n**Multi-agent system — tracking handoffs:**\n```python\nfrom openai import OpenAI\n\n# Orchestrator\norchestrator_client = OpenAI(\n    base_url=\"http://localhost:8000/v1\",\n    api_key=\"your-openai-key\",\n    default_headers={\n        \"x-agenticledger-session-id\":  \"run-abc123\",\n        \"x-agenticledger-agent-name\":  \"orchestrator\",\n    },\n)\n\n# Researcher (receives handoff from orchestrator)\nresearcher_client = OpenAI(\n    base_url=\"http://localhost:8000/v1\",\n    api_key=\"your-openai-key\",\n    default_headers={\n        \"x-agenticledger-session-id\":   \"run-abc123\",\n        \"x-agenticledger-agent-name\":   \"researcher\",\n        \"x-agenticledger-handoff-from\": \"orchestrator\",\n        \"x-agenticledger-handoff-to\":   \"researcher\",\n    },\n)\n```\n\nThe Flow tab renders `orchestrator → researcher` as a DAG with cost and latency on each node.\n\n**OpenAI Agents SDK (`openai-agents`) — per-agent clients:**\n\nThe `openai-agents` SDK uses its own internal OpenAI client. To pass Agentic Ledger headers you need to create a client per agent using `OpenAIResponsesModel` and set it as the agent's `model`.\n\n```python\nimport uuid\nimport os\nfrom openai import AsyncOpenAI\nfrom agents import Agent\nfrom agents.models.openai_responses import OpenAIResponsesModel\n\nSESSION_ID = f\"run-{uuid.uuid4().hex[:8]}\"  # one per execution\nBASE_URL = os.getenv(\"OPENAI_BASE_URL\")      # e.g. http://localhost:8000/v1\n\ndef al_model(agent_name: str, model: str = \"gpt-4o-mini\",\n             handoff_from: str | None = None, handoff_to: str | None = None):\n    \"\"\"Create a model instance that sends Agentic Ledger metadata headers.\"\"\"\n    if not BASE_URL:\n        return model  # proxy not configured — use default client\n    headers = {\n        \"x-agenticledger-session-id\": SESSION_ID,\n        \"x-agenticledger-agent-name\": agent_name,\n    }\n    if handoff_from:\n        headers[\"x-agenticledger-handoff-from\"] = handoff_from\n    if handoff_to:\n        headers[\"x-agenticledger-handoff-to\"] = handoff_to\n    client = AsyncOpenAI(base_url=BASE_URL, api_key=os.getenv(\"OPENAI_API_KEY\", \"\"),\n                         default_headers=headers)\n    return OpenAIResponsesModel(model=model, openai_client=client)\n\nplanner = Agent(name=\"PlannerAgent\", model=al_model(\"PlannerAgent\", handoff_to=\"SearchAgent\"), ...)\nsearcher = Agent(name=\"SearchAgent\",  model=al_model(\"SearchAgent\",  handoff_from=\"PlannerAgent\", handoff_to=\"WriterAgent\"), ...)\nwriter   = Agent(name=\"WriterAgent\",  model=al_model(\"WriterAgent\",  handoff_from=\"SearchAgent\",  handoff_to=\"EmailAgent\"), ...)\nemailer  = Agent(name=\"EmailAgent\",   model=al_model(\"EmailAgent\",   handoff_from=\"WriterAgent\"), ...)\n```\n\nEach agent's calls are tagged with its name and pipeline position. The Flow tab renders the full `PlannerAgent → SearchAgent → WriterAgent → EmailAgent` DAG automatically.\n\n> **Why per-agent clients?** `set_default_openai_client()` sets a single global client — fine for single-agent apps, but it can't carry different `agent_name` or `handoff_*` headers per agent in a multi-agent system. Per-agent `OpenAIResponsesModel` instances are the correct approach.\n\n---\n\n## Alerts\n\nAgentic Ledger fires a `POST` to your webhook URL when a threshold is breached. You connect it to whatever you already use — Slack, PagerDuty, Discord, email, or a custom endpoint. Agentic Ledger sends the payload; the integration is on your side.\n\n**Payload format:**\n```json\n{\n  \"type\":       \"high_cost\",\n  \"message\":    \"Single call cost $0.1842 exceeded threshold $0.10\",\n  \"value\":      0.1842,\n  \"threshold\":  0.10,\n  \"action_id\":  \"a1b2c3d4-...\",\n  \"session_id\": \"run-1\",\n  \"agent_name\": \"researcher\",\n  \"timestamp\":  \"2026-04-03T12:00:00+00:00\"\n}\n```\n\n**Alert types:**\n\n| Type | Triggered when |\n|---|---|\n| `high_cost` | A single call exceeds `AGENTICLEDGER_ALERT_COST_PER_CALL` |\n| `high_latency` | A single call takes longer than `AGENTICLEDGER_ALERT_LATENCY_MS` |\n| `high_error_rate` | Session error rate exceeds `AGENTICLEDGER_ALERT_ERROR_RATE` |\n| `daily_spend` | Daily total spend crosses `AGENTICLEDGER_ALERT_DAILY_SPEND` |\n| `budget_exceeded` | A budget limit is hit and `AGENTICLEDGER_BUDGET_ACTION` is `warn` or `both` |\n| `loop_flag` | The loop engine raised flags on a call (`repeat_tool_call`, `step_budget_exceeded`, `completion_promise`) |\n| `run_complete` | A run's completion promise was seen — the payload carries the full run summary (iterations, cost, tokens, flagged calls) |\n\n### Team cards — one proxy, many teams\n\nThink allowance cards: you keep the one real provider key, and hand each\nteam a card of its own. Each card opens the proxy, stamps every call with\nthe team's name, and can carry its own daily budget — when marketing hits\n$10, only marketing gets blocked (with an honest `Retry-After`).\n\n```bash\ncurl -X POST http://localhost:8000/api/tokens \\\n  -H \"x-agenticledger-api-key: $ADMIN_KEY\" -H 'content-type: application/json' \\\n  -d '{\"name\": \"marketing\", \"role\": \"ingest\", \"budget_daily\": 10.00}'\n```\n\nThe response shows the card once — the ledger stores only its hash. The\nteam puts it in `x-agenticledger-ingest-key` instead of the shared key;\nReports gains a by-team table with errors, blocks, and spend-today against\neach card's allowance. Revoke a card with `DELETE /api/tokens/{token_id}`\nand only that team is affected — from that instant the card gets a final\n**403** (\"the answer is no\"), which agents accept without retry storms.\nPaste a card into the dashboard's ⚿ panel by mistake and it tells you, in\nplain words, that cards open the relay, not the dashboard.\n\n**Budgets vs alerts:**\n- **Budgets** (`AGENTICLEDGER_BUDGET_*`) — block the call before it reaches the LLM. Agent gets HTTP 429.\n- **Alerts** (`AGENTICLEDGER_ALERT_*`) — the call goes through, you get notified after.\n\n**Slack** — create an [Incoming Webhook](https://api.slack.com/messaging/webhooks) and point `AGENTICLEDGER_ALERT_WEBHOOK_URL` at it. Add `AGENTICLEDGER_DIGEST_HOUR=8` and the same webhook also gets a daily good-morning digest: last-24h spend, cache savings, and the top models and agents.\n\n**PagerDuty** — use the [Events API v2](https://developer.pagerduty.com/docs/events-api-v2/) URL or a thin adapter that maps `type` → PagerDuty severity.\n\n**Discord** — use a Discord channel webhook URL directly.\n\n**Custom** — any HTTP endpoint that accepts a JSON `POST`.\n\n---\n\n## OpenTelemetry export\n\nAgentic Ledger can emit every intercepted LLM call as an OTel span to any OTLP-compatible collector: Grafana Tempo, Jaeger, Honeycomb, Datadog, Dynatrace, or any vendor that supports OTLP/HTTP.\n\n**Install the extra** (Docker image includes OTel — no extra step needed when using Docker):\n```bash\npip install \"agentic-ledger[otel]\"\n# or\nuv add \"agentic-ledger[otel]\"\n```\n\n**Configure:**\n\n| Variable | Default | Description |\n|---|---|---|\n| `AGENTICLEDGER_OTEL_ENDPOINT` | _(none)_ | OTLP/HTTP base URL, e.g. `http://localhost:4318`. OTel export is disabled when not set. |\n| `AGENTICLEDGER_OTEL_SERVICE_NAME` | `agenticledger` | Value of `service.name` in the emitted resource. |\n| `AGENTICLEDGER_OTEL_HEADERS` | _(none)_ | Comma-separated `key=value` pairs for auth headers, e.g. `x-honeycomb-team=abc123,x-honeycomb-dataset=llm`. |\n\n**Example — Grafana Tempo:**\n```bash\nAGENTICLEDGER_UPSTREAM_URL=https://api.openai.com \\\nAGENTICLEDGER_OTEL_ENDPOINT=http://localhost:4318 \\\nAGENTICLEDGER_OTEL_SERVICE_NAME=my-agent \\\nuv run python -m agenticledger.proxy\n```\n\n**Example — Honeycomb:**\n```bash\nAGENTICLEDGER_OTEL_ENDPOINT=https://api.honeycomb.io \\\nAGENTICLEDGER_OTEL_HEADERS=x-honeycomb-team=YOUR_API_KEY,x-honeycomb-dataset=llm-traces \\\nuv run python -m agenticledger.proxy\n```\n\n**Span attributes emitted (GenAI semantic conventions):**\n\n| Attribute | Source |\n|---|---|\n| `gen_ai.system` | Provider (`openai` / `anthropic`) |\n| `gen_ai.operation.name` | Always `chat` |\n| `gen_ai.request.model` | Model ID |\n| `gen_ai.request.temperature` | If set |\n| `gen_ai.request.max_tokens` | If set |\n| `gen_ai.usage.input_tokens` | Tokens in |\n| `gen_ai.usage.output_tokens` | Tokens out |\n| `gen_ai.response.finish_reasons` | Stop reason |\n| `agenticledger.action_id` | Unique call ID |\n| `agenticledger.session_id` | Run grouping |\n| `agenticledger.agent_name` | From header |\n| `agenticledger.user_id` | From header |\n| `agenticledger.cost_usd` | Estimated cost |\n| `agenticledger.latency_ms` | End-to-end latency |\n| `agenticledger.environment` | From header |\n| `agenticledger.handoff_from` / `agenticledger.handoff_to` | Agent handoffs |\n| `http.status_code` | HTTP status from upstream |\n\nSpans are grouped into traces by `session_id` — all calls in a session appear as one trace in your backend. Parent-child relationships follow `x-agenticledger-parent-action-id`. Error spans (`status_code != 200`) are marked with `StatusCode.ERROR`.\n\n---\n\n## Compliance export\n\nEvery session can be exported as an integrity-tagged audit trail — useful for regulated industries, internal audits, or passing traces to external tools.\n\n```bash\n# Machine-readable JSON with an integrity tag over the calls array\ncurl http://localhost:8000/export/run-1 -o audit-run-1.json\n\n# Printable HTML — open in browser and print to PDF\nopen http://localhost:8000/export/run-1/report\n```\n\nThe JSON export carries an integrity tag over the calls array. By default this is a `sha256` **checksum** — it catches accidental corruption but is not a signature (anyone who edits the calls can recompute it). Set `AGENTICLEDGER_EXPORT_HMAC_KEY` to switch to a keyed **`hmac-sha256`** tag, which is tamper-evident: a recipient holding the key can detect any modification, and the tag cannot be forged without the key.\n\n---\n\n## Releasing\n\nTagging a version triggers the full release pipeline automatically:\n\n```bash\ngit tag v0.2.0\ngit push origin v0.2.0\n```\n\nThis runs three jobs:\n1. **Docker** — builds `ghcr.io/shekharbhardwaj/agentic-ledger:{version}` and `:latest` for linux/amd64 + linux/arm64, pushes with SBOM + provenance attestations, signs the digest with Sigstore cosign (keyless), and mirrors to Docker Hub when the `DOCKERHUB_USERNAME`/`DOCKERHUB_TOKEN` secrets are configured\n2. **PyPI** — builds and publishes `agentic-ledger=={version}` to PyPI using trusted publishing (no API token needed), with PEP 740 attestations\n3. **GitHub Release** — creates a release with auto-generated changelog and attaches the image SBOM (SPDX)\n\n**First-time PyPI setup** (one time only):\n1. Go to [pypi.org/manage/account/publishing](https://pypi.org/manage/account/publishing/)\n2. Add a new pending publisher:\n   ```\n   PyPI project name:  agentic-ledger\n   Owner:              ShekharBhardwaj\n   Repository:         AgenticLedger\n   Workflow name:      release.yml\n   Environment name:   pypi\n   ```\n3. Create a `pypi` environment in GitHub: repo → Settings → Environments → New environment → name it `pypi`\n4. That's it — no secrets needed\n\n---\n\n## Troubleshooting\n\n**Start here: `agenticledger doctor`.** One command prints the whole truth Add `--fix` and it applies the fixes it names: shadow installs evicted with their own interpreter's pip, the PATH prepend offered for cleanup, then a second diagnostic pass.\nof your machine: every install on PATH and who shadows whom, which Python\nowns each one and whether it can actually run (wrong-architecture wheels\nand missing dependencies caught by a real import probe), what the\nbackground service is serving, and a fix-it command per finding. Most of\nthe problems below diagnose themselves with it.\n\n**Old version / commands or env vars named `agentledger` (no \"ic\")** —\nyou're running a pre-0.4 release, most likely from a venv that already had\nthe package installed: plain `pip install agentic-ledger` says \"requirement\nalready satisfied\" and does NOT upgrade. Run `agenticledger upgrade` — it\nuses the Python environment that owns the install, so there's no guessing\nwhich pip is the right one — then restart. The proxy prints its version on the first line at startup, and\n`curl localhost:8000/health` reports it too. Since 0.4.0 everything is named\n`agenticledger` — see the migration notes in the CHANGELOG.\n\n**Replay fails with 401 \"invalid x-api-key\"** — `AGENTICLEDGER_REPLAY_API_KEY`\nneeds a real provider API key from [console.anthropic.com](https://console.anthropic.com)\n(or platform.openai.com). A Claude Code subscription login is **not** an API\nkey and cannot be used. No key? Replay for free against a local model — see\nthe [LM Studio guide](docs/integrations/lm-studio.md).\n\n**`incompatible architecture (have 'arm64', need 'x86_64')`** on macOS — your\nterminal is running under Rosetta, so Python picks its x86_64 slice while pip\ninstalled arm64 native wheels. Check with `arch` (should print `arm64` on\nApple Silicon). Quick fix: prefix the command with `arch -arm64`. Permanent\nfix: uncheck \"Open using Rosetta\" on your terminal app, use an Apple Silicon\nbuild of your editor, and restart any long-lived `tmux` server.\n\n**`module 'httpx' has no attribute 'AsyncClient'`** — fixed in\n`0.3.0-alpha.2`; upgrade with `pip install --upgrade agentic-ledger`.\n\n**Port 8000 already in use** — another proxy instance (or app) is running;\nstop it or set `AGENTICLEDGER_PORT`.\n\n**`401 OAuth access token has expired` from Claude Code** — the proxy passed\nAnthropic's answer through unmodified; re-authenticate with `claude` →\n`/login`. Errored calls are still captured, so you'll see the 401 in the\ndashboard.\n\n**`/` answers 404 \"Web app not built\"** — you're running from a source\ncheckout without the web-app build. `cd dashboard-app && npm ci &&\nnpm run build` and restart. PyPI and Docker installs always include the app.\n\n---\n\n## License\n\nMIT\n\n<!-- MCP registry ownership verification -->\n`mcp-name: io.github.ShekharBhardwaj/agentic-ledger`\n",
  "bytes": 57889,
  "sha": "ad40317b9b8399a5f14c352b52dd0faace4b84067422b3273e083a4c859197fd",
  "repo_slug": "shekharbhardwaj/agenticledger",
  "fonte": "repo",
  "truncated": false,
  "api": "https://agentalog.com/api/listings/mcp_io_github_shekharbhardwaj_agentic_ledger_cad3de0d/readme"
}