Faultline
Check infrastructure health, manage incidents, and run runbooks in Faultline.
Open source Repository Open in the app JSON README (API)
About
Check infrastructure health, manage incidents, and run runbooks in Faultline.
Details
- Kind
- MCP servers
- Topic
- Cloud & DevOps
- Publisher
- io.fltln
- Origin
- official
- Category
- ferramentas
- Transport
- http
- Version
- 2.0.0
- Last push
- 2026-07-04T21:57:57Z
- Repository state
- ativo
- Added
- 2026-08-29 03:01:37
- Updated
- 2026-08-29 03:01:37
- Origin id
io.fltln/faultline
README
# Faultline MCP Server
[Faultline](https://fltln.io) is infrastructure monitoring and incident
management for DevOps/SRE teams. This repo documents `faultline-mcp` —
Faultline's remote MCP server, which lets AI agents (Claude, Claude Code,
Claude Desktop, or anything MCP-compatible) operate Faultline: check
infrastructure health, inspect and act on incidents, look up who's on call,
and run approved runbooks.
**This repo is documentation only.** The server is hosted by Faultline at
`https://mcp.fltln.io/mcp` (Streamable HTTP transport) — there's nothing to
install or run yourself.
## Setup
1. Create an API key in Faultline: **Settings → API Keys**. Keys look like
`flt_...`.
2. Point your MCP client at `https://mcp.fltln.io/mcp`, sending the key as
either `X-API-Key: flt_...` or `Authorization: Bearer flt_...`.
**Claude Code:**
```bash
claude mcp add --transport http faultline-mcp https://mcp.fltln.io/mcp \
--header "X-API-Key: flt_..."
```
**Clients that take raw JSON config** (Claude Desktop, etc.):
```json
{
"mcpServers": {
"faultline": {
"type": "http",
"url": "https://mcp.fltln.io/mcp",
"headers": { "X-API-Key": "flt_..." }
}
}
}
```
## Tools
| Tool | What it does |
| --- | --- |
| `list_services` | Monitor inventory with current status (optional status filter) |
| `get_service` | One service + its 10 most recent checks (for diagnosis) |
| `list_incidents` | Open incidents (or `status: "resolved"` for history) |
| `get_incident` | Full incident record: timeline, AI summary, post-mortem |
| `acknowledge_incident` | Acknowledge an incident — stops further escalation |
| `resolve_incident` | Resolve with an optional note (recorded on the timeline) |
| `who_is_on_call` | Current on-call per schedule, with shift end time |
| `list_anomalies` | Recent learned-baseline latency anomalies (observed vs baseline, z-score, hours sustained, auto-opened incident if any) |
| `diagnose_incident` | Recommend the next action (run runbook / escalate / resolve / wait) + candidate runbooks. Analysis only — changes nothing |
| `run_runbook` | Execute one chosen runbook against an incident — **mutates infrastructure** (can restart/scale services) |
## Security
- **Scoped to your API key.** The server never stores your key — it's used
only for the duration of each request, proxied straight through to
Faultline's API.
- **Approval-gated mutation.** `run_runbook` is the only tool that changes
infrastructure. Its description instructs the calling agent to use it only
after `diagnose_incident` recommended it and you've explicitly confirmed.
- **Tenant-isolated.** Every request is scoped to the tenant that owns the
API key — one key can never see or affect another tenant's data.
## Support
Questions or issues: [support@fltln.io](mailto:support@fltln.io) or the
[Faultline dashboard](https://fltln.io).